Qwen3.8-Flash-Next
Qwen3.8-Flash-Next by Alibaba previews the Qwen4 architecture with Qwen Sparse Attention, Gated Residual, N-gram embeddings, and 6B active parameters.
Model Overview
Capabilities, design details, and architectural traits
Qwen3.8-Flash-Next - Experimental Preview of the Qwen4 Architecture
Qwen3.8-Flash-Next is an open-weight multimodal MoE model from Alibaba's Qwen team that serves as an early preview of the architecture that will underpin Qwen4. It plays the same role Qwen3-Next played for the Qwen3.5 series: the architectural changes are released ahead of the full model family so the community can examine them first. The release upgrades the model along four axes - attention, residual, embedding, and optimization - with the stated goal of ultimate cost-efficiency.
| Trait | Detail |
|---|---|
| Qwen4 architecture preview | Open-weight release of the design that the full Qwen4 family will be built on |
| Hybrid attention | Gated DeltaNet compresses history efficiently; Qwen Sparse Attention (QSA) uses a lightweight indexer to select important context at micro-block granularity, cutting long-context latency |
| Gated Residual | Widens the residual stream into 4 branches, modulated by an element-wise data-dependent read gate and a per-branch scalar write gate |
| N-gram Embedding | Indexed by bigrams and trigrams at layer 2, offloadable to host memory and overlapped with compute through asynchronous prefetching |
| Tailored training recipe | Muon and AdamW optimizers assigned to specific weight categories, with refitted scaling laws and no batch-size warmup |
| Sparse activation | 512 experts (10 routed + 1 shared) |
| Long context | 262,144 tokens natively, extensible to 1,000,000 with YaRN |
Cost-Efficiency as the Defining Goal
The architecture is explicitly built around efficiency rather than raw scale. Compared with Qwen3.7-Plus, training takes about one-ninth the cost while the model delivers stronger results in coding and office tasks, and long-sequence prefill and decode are substantially faster.
Production Version
A managed version, Qwen3.8-Flash, is served on Qwen Cloud with 1M context length by default and official built-in tools. The open weights are compatible with Hugging Face Transformers, vLLM, and SGLang.
Benchmark Performance
Independent evaluations · Artificial Analysis
Accuracy & Capability Details
Compare Models Side-by-Side
Evaluate specifications, pricing, and independent benchmark indices
| Model Details | |||
|---|---|---|---|
| General Info | |||
| Provider | Alibaba | Anthropic | Anthropic |
| Release Date | August 26, 2026 | September 22, 2026 | September 28, 2026 |
| Knowledge Cutoff | - | - | Jun 2026 |
| Context & Limits | |||
| Context Window | 262K | 1M Best Context Window | 1M Best Context Window |
| Pricing (per 1M tokens) | |||
| Input Pricing | $0.15 Best Input Pricing | $4 | $2 |
| Output Pricing | $0.47 Best Output Pricing | $20 | $10 |
| Modalities | |||
| Inputs | textimagevideo | textimagefile | textimagefile |
| Outputs | text | text | text |
| Benchmarks (0-100) | |||
| Intelligence Index | 39.8 | 57.6 Best Intelligence Index | 56.0 |
| Coding Index | 73.1 | - | - |
| Agentic Index | 53.6 | - | - |
Humanity's Last Exam
Extremely difficult logical reasoning and knowledge.
Long Context Reasoning
Logical reasoning over long context windows.
SciCode Benchmark
Scientific coding and mathematical modeling.
Independent evaluation data provided by Artificial Analysis. To view the latest benchmarks and full details, visit their official site.
Explore more from Alibaba
Other models by Alibaba
Top AI Models
Leading alternatives by intelligence score