Google
Released September 2, 2026

Gemini 3.8 Flash

Gemini 3.8 Flash by Google is its most intelligent Flash model, built for long-horizon software engineering, autonomous agents, and enterprise workflows.

Inputs
Text
Outputs
Text

Model Overview

Capabilities, design details, and architectural traits

Gemini 3.8 Flash - Google's most intelligent Flash model for long-horizon work

Gemini 3.8 Flash is Google's most intelligent Flash model, engineered for long-horizon software engineering, autonomous agents, and complex enterprise workflows. It is generally available and ready for production use, and it delivers its gains at the same speed and low cost as the previous 3.7 Flash.

TraitDetail
PositioningMost intelligent Flash model in the Gemini lineup, often approaching higher-cost frontier models
Built forLong-horizon software engineering, autonomous agents, and complex enterprise workflows
Reasoning controlTunable thinking levels: low, medium, high, with medium as the default
Context1M token context window and 64k max output tokens
Training edgeCoding and reasoning gains driven in part by rigorous training in the demanding domain of cybersecurity

Tunable thinking levels

A defining feature of 3.8 Flash is its reasoning control. Developers can set the thinking level to low, medium, or high, trading depth against latency and cost per task. The default is medium.

A cybersecurity-trained sibling

3.8 Flash shares its foundational intelligence with Gemini 3.8 Flash Cyber, a variant focused on vulnerability detection and automated patching. Both variants are accelerated by long-running agentic loops that recursively evaluate and refine the underlying models, and the cybersecurity training behind this shared core is a documented driver of 3.8 Flash's coding and reasoning gains.

Benchmark Performance

Independent evaluations · Artificial Analysis

58.7%
Intelligence
76.3%
Coding Index
50.0%
Agentic Index

Accuracy & Capability Details

GPQA - Graduate Science95.3%
Humanity's Last Exam47.8%
SciCode - Scientific Coding53.6%
Long Context Reasoning81.0%

Compare Models Side-by-Side

Evaluate specifications, pricing, and independent benchmark indices

Model Details
General Info
ProviderGoogleAnthropicAnthropic
Release DateSeptember 2, 2026September 1, 2026July 24, 2026
Knowledge Cutoff--May 2026
Context & Limits
Context Window--1M
Pricing (per 1M tokens)
Input Pricing
$0.75
Best Input Pricing
$10$5
Output Pricing
$3.75
Best Output Pricing
$50$25
Modalities
Inputs
text
textimagefile
textimage
Outputs
text
text
text
Benchmarks (0-100)
Intelligence Index58.7
65.7
Best Intelligence Index
63.1
Coding Index76.3
81.6
Best Coding Index
78.0
Agentic Index50.0
61.3
Best Agentic Index
59.2
Gemini 3.8 Flash
Claude Fable 5.1
Claude Opus 5

GPQA Benchmark

Graduate-level reasoning and expert Q&A evaluation.

95%
Gemini 3.8 Flash
GPQA Benchmark
Score: 95%
Gemini 3.8 Flash
94%
Claude Fable 5.1
GPQA Benchmark
Score: 94%
Claude Fable 5.1
93%
Claude Opus 5
GPQA Benchmark
Score: 93%
Claude Opus 5

Humanity's Last Exam

Extremely difficult logical reasoning and knowledge.

48%
Gemini 3.8 Flash
Humanity's Last Exam
Score: 48%
Gemini 3.8 Flash
59%
Claude Fable 5.1
Humanity's Last Exam
Score: 59%
Claude Fable 5.1
55%
Claude Opus 5
Humanity's Last Exam
Score: 55%
Claude Opus 5

Long Context Reasoning

Logical reasoning over long context windows.

81%
Gemini 3.8 Flash
Long Context Reasoning
Score: 81%
Gemini 3.8 Flash
80%
Claude Fable 5.1
Long Context Reasoning
Score: 80%
Claude Fable 5.1
76%
Claude Opus 5
Long Context Reasoning
Score: 76%
Claude Opus 5

Independent evaluation data provided by Artificial Analysis. To view the latest benchmarks and full details, visit their official site.