Qwen3-VL-30B-A3B is Qwen's edge-scale MoE vision-language model with DeepStack fusion, Interleaved-MRoPE, 3D grounding, and a GUI-operating Visual Agent mode.
Capabilities, design details, and architectural traits
Qwen3-VL-30B-A3B is the smaller mixture-of-experts member of the Qwen3-VL line, built to carry the same architecture upgrades down to lighter deployment scale. It is released in Instruct and Thinking editions and keeps the same core vision-language redesign found across the Qwen3-VL family.
| Trait | Documented Behavior |
|---|---|
| DeepStack | Fuses multi-level Vision Transformer features to sharpen fine-grained image-text alignment. |
| Interleaved-MRoPE | Allocates positional embedding frequency across time, width, and height for long-horizon video reasoning. |
| Text-Timestamp Alignment | Grounds video events to precise timestamps, moving beyond the earlier T-RoPE method. |
| Visual Agent | Operates PC and mobile GUIs by recognizing on-screen elements, understanding their function, and invoking tools to finish tasks. |
| Visual Coding Boost | Converts images or video directly into Draw.io diagrams, HTML, CSS, or JS code. |
Qwen3-VL-30B-A3B represents the edge-to-cloud scaling promise stated for the series: the same DeepStack, Interleaved-MRoPE, and agentic Visual Agent behaviors run in a smaller mixture-of-experts footprint rather than the largest flagship configuration. It also carries the same 2D grounding and 3D grounding spatial reasoning documented for embodied AI use cases across the Qwen3-VL generation.
Independent evaluations · Artificial Analysis
Evaluate specifications, pricing, and independent benchmark indices
| Model Details | |||
|---|---|---|---|
| General Info | |||
| Provider | Alibaba | Anthropic | Anthropic |
| Release Date | October 3, 2025 | July 24, 2026 | June 9, 2026 |
| Knowledge Cutoff | Mar 2025 | May 2026 | - |
| Context & Limits | |||
| Context Window | 131K | 1M Best Context Window | 1M Best Context Window |
| Pricing (per 1M tokens) | |||
| Input Pricing | $0.20 Best Input Pricing | $5 | $10 |
| Output Pricing | $2.40 Best Output Pricing | $25 | $50 |
| Modalities | |||
| Inputs | textimage | textimage | textimagefile |
| Outputs | text | text | text |
| Benchmarks (0-100) | |||
| Intelligence Index | 13.4 | 63.1 Best Intelligence Index | 62.1 |
| Coding Index | - | 78.0 Best Coding Index | 76.5 |
| Agentic Index | - | 59.2 Best Agentic Index | 56.6 |
Graduate-level reasoning and expert Q&A evaluation.
Extremely difficult logical reasoning and knowledge.
Logical reasoning over long context windows.
Independent evaluation data provided by Artificial Analysis. To view the latest benchmarks and full details, visit their official site.