Toolbit.aiToolbit.ai

Find, compare, and explore the best AI tools to match your specific tasks and use cases.

Explore

  • AI Search
  • Compare ToolsNew
  • Browse Categories
  • Trending Tools
  • Most Popular
  • New Additions

Resources

  • Updates HubNew
  • AI News
  • ModelsNew
  • Blog Articles
  • NewsletterNew

Company

  • Launch a Tool
  • Advertise with Us
  • Guest Post
  • Contact Us
© 2026 Toolbit.ai. All rights reserved.
Privacy PolicyTerms & ConditionsDisclaimer
There's An AI For That favicon
There's An AI For That•The front page of AI for everyone
Toolbit.ai
Toolbit.ai
UpdatesNew
Blog
Toolbit.ai
Toolbit.ai
Toolbit.ai
Toolbit.ai
UpdatesNew
Blog
Sign in
  1. Home
  2. Updates
  3. Models
  4. Qwen3.5-35B-A3B
Alibaba
Released February 24, 2026

Qwen3.5-35B-A3B

Qwen3.5 35B A3B is Qwen's reasoning vision-language model, built on Gated Delta Networks and Mixture-of-Experts, with toggleable thinking and tool use.

Visit AlibabaAnnouncement
Inputs
Text
Image
Video
Outputs
Text

Model Overview

Capabilities, design details, and architectural traits

Qwen3.5 35B A3B - A Reasoning Vision-Language Model With Tool Use

Qwen3.5 35B A3B is described by the Qwen team as a reasoning vision-language model that supports tool use. It combines multimodal understanding with agentic capability in one model, built on an architecture that favors efficiency over raw size.

TraitDetail
Defining purposeA reasoning vision-language model that supports tool use, combining multimodal understanding with agentic capability in one model.
Gated Delta plus MoEBuilt on a hybrid architecture combining Gated Delta Networks with a sparse Mixture-of-Experts design, aimed at high-throughput inference with low latency.
Early fusion trainingTrained with early fusion on multimodal tokens, a unified approach to vision-language learning rather than attaching a separate vision encoder to a text only model.
Toggleable thinking modeSupports both a thinking mode for step by step reasoning and a non-thinking mode for direct responses, switchable depending on the task.

Trained With Massive-Scale Agent Reinforcement Learning

Qwen states it scaled reinforcement learning across environments involving large numbers of agents, using progressively complex task distributions aimed at robust real-world adaptability rather than a single fixed training scenario.

Benchmark Performance

Independent evaluations · Artificial Analysis

29.9%
Intelligence

Accuracy & Capability Details

GPQA - Graduate Science84.5%
Humanity's Last Exam21.0%
SciCode - Scientific Coding37.7%
Instruction Following72.5%
Long Context Reasoning68.3%
τ²-Bench - Agentic Tasks89.2%
TerminalBench - System Control26.5%
Specs
Context window
262Ktokens
Input pricing
$0.25per 1M tokens
Output pricing
$2per 1M tokens

Prices in USD.

Key Capabilities & Ratings

Scientific Biology90%
Translation & Lang90%
Structured Outputs90%
Advanced Math80%
Video80%

Compare Models Side-by-Side

Evaluate specifications, pricing, and independent benchmark indices

Model Details
General Info
ProviderAlibabaAnthropicAnthropic
Release DateFebruary 24, 2026July 24, 2026June 9, 2026
Knowledge Cutoff-May 2026-
Context & Limits
Context Window262K
1M
Best Context Window
1M
Best Context Window
Pricing (per 1M tokens)
Input Pricing
$0.25
Best Input Pricing
$5$10
Output Pricing
$2
Best Output Pricing
$25$50
Modalities
Inputs
textimagevideo
textimage
textimagefile
Outputs
text
text
text
Benchmarks (0-100)
Intelligence Index24.3
63.1
Best Intelligence Index
62.1
Coding Index37.0
78.0
Best Coding Index
76.5
Agentic Index11.8
59.2
Best Agentic Index
56.6
Qwen3.5-35B-A3B
Claude Opus 5
Claude Fable 5

GPQA Benchmark

Graduate-level reasoning and expert Q&A evaluation.

82%
Qwen3.5-35B-A3B
GPQA Benchmark
Score: 82%
Qwen3.5-35B-A3B
93%
Claude Opus 5
GPQA Benchmark
Score: 93%
Claude Opus 5
93%
Claude Fable 5
GPQA Benchmark
Score: 93%
Claude Fable 5

Humanity's Last Exam

Extremely difficult logical reasoning and knowledge.

13%
Qwen3.5-35B-A3B
Humanity's Last Exam
Score: 13%
Qwen3.5-35B-A3B
55%
Claude Opus 5
Humanity's Last Exam
Score: 55%
Claude Opus 5
56%
Claude Fable 5
Humanity's Last Exam
Score: 56%
Claude Fable 5

Long Context Reasoning

Logical reasoning over long context windows.

60%
Qwen3.5-35B-A3B
Long Context Reasoning
Score: 60%
Qwen3.5-35B-A3B
76%
Claude Opus 5
Long Context Reasoning
Score: 76%
Claude Opus 5
77%
Claude Fable 5
Long Context Reasoning
Score: 77%
Claude Fable 5

Independent evaluation data provided by Artificial Analysis. To view the latest benchmarks and full details, visit their official site.

Back to all models