Qwen3-VL-235B-A22B is Qwen's MoE vision-language model with DeepStack fusion, Interleaved-MRoPE reasoning, 3D grounding, and a GUI-operating Visual Agent mode.
Capabilities, design details, and architectural traits
Qwen3-VL-235B-A22B is Qwen's mixture-of-experts vision-language model, released in Instruct and Thinking editions. It combines a redesigned vision-language architecture with agentic features built for operating software interfaces and turning visual input directly into working code.
| Trait | Documented Behavior |
|---|---|
| DeepStack | Fuses multi-level Vision Transformer features to sharpen fine-grained image-text alignment. |
| Interleaved-MRoPE | Allocates positional embedding frequency across time, width, and height for long-horizon video reasoning. |
| Text-Timestamp Alignment | Grounds video events to precise timestamps, moving beyond the earlier T-RoPE method. |
| Visual Agent | Operates PC and mobile GUIs by recognizing on-screen elements, understanding their function, and invoking tools to finish tasks. |
| Visual Coding Boost | Converts images or video directly into Draw.io diagrams, HTML, CSS, or JS code. |
Qwen3-VL-235B-A22B extends 2D grounding into 3D grounding, letting it judge object position, viewpoint, and occlusion for embodied AI tasks. Its OCR pipeline covers 32 languages and is documented as robust to low light, blur, and tilt, with improved handling of rare and ancient characters.
Independent evaluations · Artificial Analysis
Evaluate specifications, pricing, and independent benchmark indices
| Model Details | |||
|---|---|---|---|
| General Info | |||
| Provider | Alibaba | Anthropic | Anthropic |
| Release Date | September 23, 2025 | July 24, 2026 | June 9, 2026 |
| Knowledge Cutoff | Mar 2025 | May 2026 | - |
| Context & Limits | |||
| Context Window | 131K | 1M Best Context Window | 1M Best Context Window |
| Pricing (per 1M tokens) | |||
| Input Pricing | $0.40 Best Input Pricing | $5 | $10 |
| Output Pricing | $1.60 Best Output Pricing | $25 | $50 |
| Modalities | |||
| Inputs | textimage | textimage | textimagefile |
| Outputs | text | text | text |
| Benchmarks (0-100) | |||
| Intelligence Index | 14.4 | 63.1 Best Intelligence Index | 62.1 |
| Coding Index | - | 78.0 Best Coding Index | 76.5 |
| Agentic Index | - | 59.2 Best Agentic Index | 56.6 |
Graduate-level reasoning and expert Q&A evaluation.
Extremely difficult logical reasoning and knowledge.
Logical reasoning over long context windows.
Independent evaluation data provided by Artificial Analysis. To view the latest benchmarks and full details, visit their official site.