Toolbit.aiToolbit.ai

Find, compare, and explore the best AI tools to match your specific tasks and use cases.

Explore

  • AI Search
  • Compare ToolsNew
  • Browse Categories
  • Trending Tools
  • Most Popular
  • New Additions

Resources

  • Updates HubNew
  • AI News
  • ModelsNew
  • Blog Articles
  • NewsletterNew

Company

  • Launch a Tool
  • Advertise with Us
  • Guest Post
  • Contact Us
© 2026 Toolbit.ai. All rights reserved.
Privacy PolicyTerms & ConditionsDisclaimer
There's An AI For That favicon
There's An AI For That•The front page of AI for everyone
Toolbit.ai
Toolbit.ai
UpdatesNew
Blog
Toolbit.ai
Toolbit.ai
Toolbit.ai
Toolbit.ai
UpdatesNew
Blog
Sign in
  1. Home
  2. Updates
  3. Models
  4. Qwen3.6 Max Preview
Alibaba
Released April 20, 2026

Qwen3.6 Max Preview

Qwen3.6 Max Preview by Alibaba is an early preview proprietary model with improved agentic coding, stronger world knowledge, preserve_thinking API, and active iteration.

Visit AlibabaAnnouncement
Inputs
Text
Outputs
Text

Model Overview

Capabilities, design details, and architectural traits

Qwen3.6 Max Preview - "Smarter, Sharper, Still Evolving"

Qwen3.6 Max Preview is officially described as an early preview of Alibaba's next proprietary model - explicitly labeled as still under active development with further improvements expected. It sits above Qwen3.6-Plus in the 3.6 family and is text-only (no image or video input), unlike the multimodal Qwen3.6-Plus it succeeds. It is available via Alibaba Cloud Model Studio as qwen3.6-max-preview and is not open-weight.

Three Documented Capability Improvements Over Qwen3.6-Plus

TraitDetail
Agentic coding gainsOfficially documented improvements across SkillsBench, SciCode, NL2Repo, and Terminal-Bench 2.0; text-only model focused on coding agent execution
Stronger world knowledge and instruction followingDocumented gains in SuperGPQA, QwenChineseBench, and ToolcallFormatIFBench over Qwen3.6-Plus
preserve_thinking featureCarries <think> reasoning blocks across all preceding turns; explicitly recommended for agentic tasks
OpenAI and Anthropic API compatibilitySupported across Beijing, Singapore, and US-Virginia regional endpoints

The "preview" status is central to this model's documented identity - Alibaba explicitly states it is iterating on the model and welcomes community feedback before a stable release.

Benchmark Performance

Independent evaluations · Artificial Analysis

41.1%
Intelligence

Accuracy & Capability Details

GPQA - Graduate Science88.8%
Humanity's Last Exam30.8%
SciCode - Scientific Coding46.9%
Instruction Following76.6%
Long Context Reasoning72.0%
τ²-Bench - Agentic Tasks95.9%
TerminalBench - System Control43.9%
Specs
Context window
262Ktokens
Input pricing
$1.30per 1M tokens
Output pricing
$7.80per 1M tokens
Cached input
$0.13per 1M tokens

Prices in USD.

Compare Models Side-by-Side

Evaluate specifications, pricing, and independent benchmark indices

Model Details
General Info
ProviderAlibabaAnthropicAnthropic
Release DateApril 20, 2026July 24, 2026June 9, 2026
Knowledge Cutoff-May 2026-
Context & Limits
Context Window262K
1M
Best Context Window
1M
Best Context Window
Pricing (per 1M tokens)
Input Pricing
$1.30
Best Input Pricing
$5$10
Output Pricing
$7.80
Best Output Pricing
$25$50
Modalities
Inputs
text
textimage
textimagefile
Outputs
text
text
text
Benchmarks (0-100)
Intelligence Index41.1
63.1
Best Intelligence Index
62.1
Coding Index-
78.0
Best Coding Index
76.5
Agentic Index-
59.2
Best Agentic Index
56.6
Qwen3.6 Max Preview
Claude Opus 5
Claude Fable 5

GPQA Benchmark

Graduate-level reasoning and expert Q&A evaluation.

89%
Qwen3.6 Max Preview
GPQA Benchmark
Score: 89%
Qwen3.6 Max Preview
93%
Claude Opus 5
GPQA Benchmark
Score: 93%
Claude Opus 5
93%
Claude Fable 5
GPQA Benchmark
Score: 93%
Claude Fable 5

Humanity's Last Exam

Extremely difficult logical reasoning and knowledge.

31%
Qwen3.6 Max Preview
Humanity's Last Exam
Score: 31%
Qwen3.6 Max Preview
55%
Claude Opus 5
Humanity's Last Exam
Score: 55%
Claude Opus 5
56%
Claude Fable 5
Humanity's Last Exam
Score: 56%
Claude Fable 5

Long Context Reasoning

Logical reasoning over long context windows.

72%
Qwen3.6 Max Preview
Long Context Reasoning
Score: 72%
Qwen3.6 Max Preview
76%
Claude Opus 5
Long Context Reasoning
Score: 76%
Claude Opus 5
77%
Claude Fable 5
Long Context Reasoning
Score: 77%
Claude Fable 5

Independent evaluation data provided by Artificial Analysis. To view the latest benchmarks and full details, visit their official site.

Back to all models