Toolbit.aiToolbit.ai

Find, compare, and explore the best AI tools to match your specific tasks and use cases.

Explore

  • AI Search
  • Compare ToolsNew
  • Browse Categories
  • Trending Tools
  • Most Popular
  • New Additions

Resources

  • Updates HubNew
  • AI News
  • ModelsNew
  • Blog Articles
  • NewsletterNew

Company

  • Launch a Tool
  • Advertise with Us
  • Guest Post
  • Contact Us
© 2026 Toolbit.ai. All rights reserved.
Privacy PolicyTerms & ConditionsDisclaimer
There's An AI For That favicon
There's An AI For That•The front page of AI for everyone
Toolbit.ai
Toolbit.ai
UpdatesNew
Blog
Toolbit.ai
Toolbit.ai
Toolbit.ai
Toolbit.ai
UpdatesNew
Blog
Sign in
  1. Home
  2. Updates
  3. Models
  4. Qwen3.8 Max
Alibaba
Released August 3, 2026

Qwen3.8 Max

Qwen3.8 Max: 2.4T sparse MoE, 95B active, 1M context. Excels in autonomous coding, research reproduction, and multimodal understanding of documents.

Visit AlibabaAnnouncement
Inputs
Text
Image
File
Video
Outputs
Text

Model Overview

Capabilities, design details, and architectural traits

Qwen3.8-Max - First Open-Weight Max-Class Model with Long-Horizon Autonomous Coding

Qwen3.8-Max is Alibaba's most capable Qwen model to date and the first Max-class model the company has committed to open-sourcing. Built on the Qwen 3.5 architecture, it combines a Sparse Mixture-of-Experts design with a hybrid attention mechanism, totaling 2.4 trillion parameters while activating only 95 billion per query.

TraitDetail
Sparse MoE at frontier scale2.4T total parameters, 95B active, balancing massive capacity with inference efficiency
1-million-token contextIngests hundred-page documents, full television series, or 100-hour livestreams into searchable knowledge bases
Self-evolving autonomous codingRan a 16-day fully autonomous coding project, building the open-sourced oh-my-cli harness with 265 commits and 127 PRs - no human intervention
Research reproduction and improvementStarting from only a paper, reproduced six findings then engineered novel methodologies that outperformed the original work over ~5 days
First open-weight Max-class releaseModel weights scheduled for public release, a first for Alibaba's Max-tier models

Feedback-Loop Self-Evolution

A defining behavior of Qwen3.8-Max is its reliance on self-evolving feedback loops rather than fixed plans. In coding, it builds an execution loop that normalizes requirements into issues, dispatches them through a state machine, runs tests, and routes failures back for re-verification. In research, it iterates experiment after experiment. In competitions, it climbs leaderboards submission after submission.

Multimodal Knowledge Construction

The model transforms static visual inputs - documents, video, screenshots - into interactive knowledge structures. It can rebuild software applications from screenshots and convert 2D architectural floor plans into 3D renderings, extending multimodal understanding beyond passive recognition into active construction.

Benchmark Performance

Independent evaluations · Artificial Analysis

58.1%
Intelligence
71.8%
Coding Index
58.4%
Agentic Index

Accuracy & Capability Details

GPQA - Graduate Science92.7%
Humanity's Last Exam43.0%
SciCode - Scientific Coding52.9%
Long Context Reasoning74.3%
Specs
Context window
1Mtokens
Input pricing
$2per 1M tokens
Output pricing
$6per 1M tokens
Cached input
$0.25per 1M tokens

Prices in USD.

Compare Models Side-by-Side

Evaluate specifications, pricing, and independent benchmark indices

Model Details
General Info
ProviderAlibabaAnthropicAnthropic
Release DateAugust 3, 2026July 24, 2026June 9, 2026
Knowledge Cutoff-May 2026-
Context & Limits
Context Window1M1M1M
Pricing (per 1M tokens)
Input Pricing
$2
Best Input Pricing
$5$10
Output Pricing
$6
Best Output Pricing
$25$50
Modalities
Inputs
textimagefilevideo
textimage
textimagefile
Outputs
text
text
text
Benchmarks (0-100)
Intelligence Index58.1
63.1
Best Intelligence Index
62.1
Coding Index71.8
78.0
Best Coding Index
76.5
Agentic Index58.4
59.2
Best Agentic Index
56.6
Qwen3.8 Max
Claude Opus 5
Claude Fable 5

GPQA Benchmark

Graduate-level reasoning and expert Q&A evaluation.

93%
Qwen3.8 Max
GPQA Benchmark
Score: 93%
Qwen3.8 Max
93%
Claude Opus 5
GPQA Benchmark
Score: 93%
Claude Opus 5
93%
Claude Fable 5
GPQA Benchmark
Score: 93%
Claude Fable 5

Humanity's Last Exam

Extremely difficult logical reasoning and knowledge.

43%
Qwen3.8 Max
Humanity's Last Exam
Score: 43%
Qwen3.8 Max
55%
Claude Opus 5
Humanity's Last Exam
Score: 55%
Claude Opus 5
56%
Claude Fable 5
Humanity's Last Exam
Score: 56%
Claude Fable 5

Long Context Reasoning

Logical reasoning over long context windows.

74%
Qwen3.8 Max
Long Context Reasoning
Score: 74%
Qwen3.8 Max
76%
Claude Opus 5
Long Context Reasoning
Score: 76%
Claude Opus 5
77%
Claude Fable 5
Long Context Reasoning
Score: 77%
Claude Fable 5

Independent evaluation data provided by Artificial Analysis. To view the latest benchmarks and full details, visit their official site.

Back to all models