Toolbit.aiToolbit.ai

Find, compare, and explore the best AI tools to match your specific tasks and use cases.

Explore

  • AI Search
  • Compare ToolsNew
  • Browse Categories
  • Trending Tools
  • Most Popular
  • New Additions

Resources

  • Updates HubNew
  • AI News
  • ModelsNew
  • Blog Articles
  • NewsletterNew

Company

  • Launch a Tool
  • Advertise with Us
  • Guest Post
  • Contact Us
© 2026 Toolbit.ai. All rights reserved.
Privacy PolicyTerms & ConditionsDisclaimer
There's An AI For That favicon
There's An AI For That•The front page of AI for everyone
Toolbit.ai
Toolbit.ai
UpdatesNew
Blog
Toolbit.ai
Toolbit.ai
Toolbit.ai
Toolbit.ai
UpdatesNew
Blog
Sign in
  1. Home
  2. Updates
  3. Models
  4. Gemini 2.5 Flash
Google
Released May 20, 2025Cutoff January 2025

Gemini 2.5 Flash

Gemini 2.5 Flash by Google DeepMind is Google's first fully hybrid reasoning model with switchable thinking, configurable thinking budgets, and native image and audio output.

Visit GoogleAnnouncement
Inputs
File
Image
Text
Audio
Video
Outputs
Text

Model Overview

Capabilities, design details, and architectural traits

Gemini 2.5 Flash - Google's First Fully Hybrid Reasoning Model

Gemini 2.5 Flash is explicitly described in its model card as Google's first fully hybrid reasoning model. This means thinking can be switched on or off per request, and developers can set thinking budgets to control the tradeoff between output quality, cost, and latency. Its architecture is sparse mixture-of-experts (MoE), which activates only a subset of parameters per input token, decoupling total model capacity from per-token serving cost.

Native Output Modalities That Extend Beyond Text

TraitDetail
Hybrid reasoning with switchable thinkingThinking can be enabled or disabled; thinking budgets let developers tune quality-cost-latency balance per request
Gemini 2.5 Flash ImageA native image output variant within the same model; supports text-to-image generation, prompt-based image editing, multi-image fusion, and character/style consistency
Gemini 2.5 Flash AudioA native audio output variant; supports live conversational agents, complex conversational workflows, and live speech-to-speech translation
1M-token context windowAccepts text, images, audio, and video as input; text output up to 64K tokens

Intended Use Position Within the 2.5 Family

The model card explicitly names three intended use cases for Gemini 2.5 Flash: cost-efficient thinking, well-rounded capabilities, and agentic tool use. This positions it between Flash-Lite (high-volume, low-latency tasks) and 2.5 Pro (maximum reasoning depth), with the hybrid thinking toggle as the primary mechanism for developers to adjust where on that spectrum each request lands.

Benchmark Performance

Independent evaluations · Artificial Analysis

20.3%
Intelligence

Accuracy & Capability Details

GPQA - Graduate Science79.0%
Humanity's Last Exam12.1%
SciCode - Scientific Coding39.4%
Instruction Following50.3%
Long Context Reasoning65.7%
τ²-Bench - Agentic Tasks31.6%
TerminalBench - System Control13.6%
Specs
Context window
1.0Mtokens
Input pricing
$0.30per 1M tokens
Output pricing
$2.50per 1M tokens
Cached input
$0.03per 1M tokens

Prices in USD.

Key Capabilities & Ratings

Translation & Lang90%
Factuality & Retrieval90%
Scientific Biology80%
Physics Problems80%
Chemistry Concepts80%

Compare Models Side-by-Side

Evaluate specifications, pricing, and independent benchmark indices

Model Details
General Info
ProviderGoogleAnthropicAnthropic
Release DateMay 20, 2025July 24, 2026June 9, 2026
Knowledge CutoffJan 2025May 2026-
Context & Limits
Context Window
1.0M
Best Context Window
1M1M
Pricing (per 1M tokens)
Input Pricing
$0.30
Best Input Pricing
$5$10
Output Pricing
$2.50
Best Output Pricing
$25$50
Modalities
Inputs
fileimagetextaudiovideo
textimage
textimagefile
Outputs
text
text
text
Benchmarks (0-100)
Intelligence Index14.2
63.1
Best Intelligence Index
62.1
Coding Index-
78.0
Best Coding Index
76.5
Agentic Index-
59.2
Best Agentic Index
56.6
Gemini 2.5 Flash
Claude Opus 5
Claude Fable 5

GPQA Benchmark

Graduate-level reasoning and expert Q&A evaluation.

68%
Gemini 2.5 Flash
GPQA Benchmark
Score: 68%
Gemini 2.5 Flash
93%
Claude Opus 5
GPQA Benchmark
Score: 93%
Claude Opus 5
93%
Claude Fable 5
GPQA Benchmark
Score: 93%
Claude Fable 5

Humanity's Last Exam

Extremely difficult logical reasoning and knowledge.

5%
Gemini 2.5 Flash
Humanity's Last Exam
Score: 5%
Gemini 2.5 Flash
55%
Claude Opus 5
Humanity's Last Exam
Score: 55%
Claude Opus 5
56%
Claude Fable 5
Humanity's Last Exam
Score: 56%
Claude Fable 5

Long Context Reasoning

Logical reasoning over long context windows.

48%
Gemini 2.5 Flash
Long Context Reasoning
Score: 48%
Gemini 2.5 Flash
76%
Claude Opus 5
Long Context Reasoning
Score: 76%
Claude Opus 5
77%
Claude Fable 5
Long Context Reasoning
Score: 77%
Claude Fable 5

Independent evaluation data provided by Artificial Analysis. To view the latest benchmarks and full details, visit their official site.

Back to all models