Toolbit.aiToolbit.ai

Find, compare, and explore the best AI tools to match your specific tasks and use cases.

Explore

  • AI Search
  • Compare ToolsNew
  • Browse Categories
  • Trending Tools
  • Most Popular
  • New Additions

Resources

  • Updates HubNew
  • AI News
  • ModelsNew
  • Blog Articles
  • NewsletterNew

Company

  • Launch a Tool
  • Advertise with Us
  • Guest Post
  • Contact Us
© 2026 Toolbit.ai. All rights reserved.
Privacy PolicyTerms & ConditionsDisclaimer
There's An AI For That favicon
There's An AI For That•The front page of AI for everyone
Toolbit.ai
Toolbit.ai
UpdatesNew
Blog
Toolbit.ai
Toolbit.ai
Toolbit.ai
Toolbit.ai
UpdatesNew
Blog
Sign in
  1. Home
  2. Updates
  3. Models
  4. Qwen3 Coder 480B A35B
Alibaba
Released July 22, 2025Cutoff June 2025

Qwen3 Coder 480B A35B

Qwen3 Coder 480B A35B is Alibaba's largest coding-specialized MoE model with 480B total and 35B active parameters, native 256K token context extended to 1M.

Visit AlibabaAnnouncement
Inputs
Text
Outputs
Text

Model Overview

Capabilities, design details, and architectural traits

Qwen3 Coder 480B A35B - Qwen's Largest Agentic Coding Model

Qwen3 Coder 480B A35B-Instruct is Qwen's most powerful agentic code model, designed for autonomous software engineering, repository-level understanding, and multi-turn tool use with interactive reasoning.

What Makes It Different: Agentic Coding with Massive Context Window

Qwen3 Coder 480B A35B sets new standards for open models in agentic coding and browser-use tasks:

Distinctive TraitWhy It MattersHow It Works
35B activated parametersEfficient inference for massive 480B modelUses 8 experts out of 160 total during inference, activating only 35B of 480B per forward pass
Native 256K token contextRepository-scale understanding in single passProcesses 262,144 tokens natively, extended to 1M tokens with Yarn extrapolation for vast codebases
Specially designed function call formatAgentic coding requires precise tool invocationSupports Qwen Code, CLINE, and other platforms with dedicated function call structure for tool use
Non-thinking mode onlyConsistent behavior for coding tasksDoes not generate blocks, spec enable_thinking=False not required for predictable output
State-of-the-art agentic performanceOpen models rival proprietary alternativesAchieves results comparable to Claude Sonnet 4 on Agentic Coding, Agentic Browser-Use, and Agentic Tool-Use benchmarks

Why It Exists

Qwen3 Coder 480B A35B was built to deliver autonomous coding agent capabilities that manage complex workflows beyond basic code generation, supporting interactive multi-turn reasoning and tool use across diverse programming ecosystems.

Benchmark Performance

Independent evaluations · Artificial Analysis

18.2%
Intelligence

Accuracy & Capability Details

GPQA - Graduate Science61.8%
Humanity's Last Exam4.5%
SciCode - Scientific Coding35.9%
Instruction Following40.5%
Long Context Reasoning45.3%
τ²-Bench - Agentic Tasks43.6%
TerminalBench - System Control18.9%
Specs
Context window
1.0Mtokens
Input pricing
$1.50per 1M tokens
Output pricing
$7.50per 1M tokens

Prices in USD.

Key Capabilities & Ratings

Financial Analysis70%
General Knowledge70%
Communication70%
Structured Outputs70%
Logical Logic60%

Compare Models Side-by-Side

Evaluate specifications, pricing, and independent benchmark indices

Model Details
General Info
ProviderAlibabaAnthropicAnthropic
Release DateJuly 22, 2025July 24, 2026June 9, 2026
Knowledge CutoffJun 2025May 2026-
Context & Limits
Context Window
1.0M
Best Context Window
1M1M
Pricing (per 1M tokens)
Input Pricing
$1.50
Best Input Pricing
$5$10
Output Pricing
$7.50
Best Output Pricing
$25$50
Modalities
Inputs
text
textimage
textimagefile
Outputs
text
text
text
Benchmarks (0-100)
Intelligence Index18.2
63.1
Best Intelligence Index
62.1
Coding Index-
78.0
Best Coding Index
76.5
Agentic Index-
59.2
Best Agentic Index
56.6
Qwen3 Coder 480B A35B
Claude Opus 5
Claude Fable 5

GPQA Benchmark

Graduate-level reasoning and expert Q&A evaluation.

62%
Qwen3 Coder 480B A35B
GPQA Benchmark
Score: 62%
Qwen3 Coder 480B A35B
93%
Claude Opus 5
GPQA Benchmark
Score: 93%
Claude Opus 5
93%
Claude Fable 5
GPQA Benchmark
Score: 93%
Claude Fable 5

Humanity's Last Exam

Extremely difficult logical reasoning and knowledge.

5%
Qwen3 Coder 480B A35B
Humanity's Last Exam
Score: 5%
Qwen3 Coder 480B A35B
55%
Claude Opus 5
Humanity's Last Exam
Score: 55%
Claude Opus 5
56%
Claude Fable 5
Humanity's Last Exam
Score: 56%
Claude Fable 5

Long Context Reasoning

Logical reasoning over long context windows.

45%
Qwen3 Coder 480B A35B
Long Context Reasoning
Score: 45%
Qwen3 Coder 480B A35B
76%
Claude Opus 5
Long Context Reasoning
Score: 76%
Claude Opus 5
77%
Claude Fable 5
Long Context Reasoning
Score: 77%
Claude Fable 5

Independent evaluation data provided by Artificial Analysis. To view the latest benchmarks and full details, visit their official site.

Back to all models