Alibaba
Released August 26, 2026

Qwen3.8-Flash-Next

Qwen3.8-Flash-Next by Alibaba previews the Qwen4 architecture with Qwen Sparse Attention, Gated Residual, N-gram embeddings, and 6B active parameters.

Inputs
Text
Image
Video
Outputs
Text

Model Overview

Capabilities, design details, and architectural traits

Qwen3.8-Flash-Next - Experimental Preview of the Qwen4 Architecture

Qwen3.8-Flash-Next is an open-weight multimodal MoE model from Alibaba's Qwen team that serves as an early preview of the architecture that will underpin Qwen4. It plays the same role Qwen3-Next played for the Qwen3.5 series: the architectural changes are released ahead of the full model family so the community can examine them first. The release upgrades the model along four axes - attention, residual, embedding, and optimization - with the stated goal of ultimate cost-efficiency.

TraitDetail
Qwen4 architecture previewOpen-weight release of the design that the full Qwen4 family will be built on
Hybrid attentionGated DeltaNet compresses history efficiently; Qwen Sparse Attention (QSA) uses a lightweight indexer to select important context at micro-block granularity, cutting long-context latency
Gated ResidualWidens the residual stream into 4 branches, modulated by an element-wise data-dependent read gate and a per-branch scalar write gate
N-gram EmbeddingIndexed by bigrams and trigrams at layer 2, offloadable to host memory and overlapped with compute through asynchronous prefetching
Tailored training recipeMuon and AdamW optimizers assigned to specific weight categories, with refitted scaling laws and no batch-size warmup
Sparse activation512 experts (10 routed + 1 shared)
Long context262,144 tokens natively, extensible to 1,000,000 with YaRN

Cost-Efficiency as the Defining Goal

The architecture is explicitly built around efficiency rather than raw scale. Compared with Qwen3.7-Plus, training takes about one-ninth the cost while the model delivers stronger results in coding and office tasks, and long-sequence prefill and decode are substantially faster.

Production Version

A managed version, Qwen3.8-Flash, is served on Qwen Cloud with 1M context length by default and official built-in tools. The open weights are compatible with Hugging Face Transformers, vLLM, and SGLang.

Benchmark Performance

Independent evaluations · Artificial Analysis

39.8%
Intelligence
73.1%
Coding Index
53.6%
Agentic Index

Accuracy & Capability Details

GPQA - Graduate Science92.3%
Humanity's Last Exam38.0%
SciCode - Scientific Coding50.6%
Long Context Reasoning79.7%

Compare Models Side-by-Side

Evaluate specifications, pricing, and independent benchmark indices

Model Details
General Info
ProviderAlibabaAnthropicAnthropic
Release DateAugust 26, 2026September 22, 2026September 28, 2026
Knowledge Cutoff--Jun 2026
Context & Limits
Context Window262K
1M
Best Context Window
1M
Best Context Window
Pricing (per 1M tokens)
Input Pricing
$0.15
Best Input Pricing
$4$2
Output Pricing
$0.47
Best Output Pricing
$20$10
Modalities
Inputs
textimagevideo
textimagefile
textimagefile
Outputs
text
text
text
Benchmarks (0-100)
Intelligence Index39.8
57.6
Best Intelligence Index
56.0
Coding Index73.1--
Agentic Index53.6--
Qwen3.8-Flash-Next
Claude Opus 5.5
Claude Sonnet 5.5

Humanity's Last Exam

Extremely difficult logical reasoning and knowledge.

38%
Qwen3.8-Flash-Next
Humanity's Last Exam
Score: 38%
Qwen3.8-Flash-Next
61%
Claude Opus 5.5
Humanity's Last Exam
Score: 61%
Claude Opus 5.5
55%
Claude Sonnet 5.5
Humanity's Last Exam
Score: 55%
Claude Sonnet 5.5

Long Context Reasoning

Logical reasoning over long context windows.

80%
Qwen3.8-Flash-Next
Long Context Reasoning
Score: 80%
Qwen3.8-Flash-Next
85%
Claude Opus 5.5
Long Context Reasoning
Score: 85%
Claude Opus 5.5
83%
Claude Sonnet 5.5
Long Context Reasoning
Score: 83%
Claude Sonnet 5.5

SciCode Benchmark

Scientific coding and mathematical modeling.

51%
Qwen3.8-Flash-Next
SciCode Benchmark
Score: 51%
Qwen3.8-Flash-Next
67%
Claude Opus 5.5
SciCode Benchmark
Score: 67%
Claude Opus 5.5
61%
Claude Sonnet 5.5
SciCode Benchmark
Score: 61%
Claude Sonnet 5.5

Independent evaluation data provided by Artificial Analysis. To view the latest benchmarks and full details, visit their official site.

Explore more from Alibaba

Other models by Alibaba

Top AI Models

Leading alternatives by intelligence score

View all