Sapiens AI
Released September 11, 2026

Agnes 3.0 Flash

Agnes 3.0 Flash by Sapiens AI is a text model for agentic coding and tool-driven tasks, with 512K context, multi-step tool orchestration, and verified

Inputs
Text
Image
Outputs
Text

Model Overview

Capabilities, design details, and architectural traits

Agnes 3.0 Flash - Agentic coding and tool-driven execution model

Agnes 3.0 Flash is Sapiens AI's next-generation text model built for agentic coding and tool-driven tasks. Its defining focus is end-to-end execution quality: the full path from task understanding and planning, through tool use, to final delivery. The model is designed to make agent applications more reliable in complex, long-running work.

DifferentiatorWhat it means
Agentic coding focusTuned for Agnes Code and coding-agent workflows, strengthening execution from requirements understanding through final delivery.
Stable tool orchestrationInterprets tool definitions, selects appropriate tools, and coordinates multi-step calls while reducing ineffective calls, repetition, and abnormal loops.
Instruction and context adherenceKeeps objectives, constraints, and runtime context across long-running, multi-turn tasks, reducing task drift and missed requirements.
Trustworthy deliveryEmphasizes factual grounding and result verification to cut unsupported conclusions and incorrect completion claims.
Clean outputReduced repetition, malformed text, and unnecessary exposure of internal reasoning, producing responses ready for delivery.
512K context windowSupports long agent sessions with up to 65,536 output tokens.
Text and image-URL inputAccepts public image URLs alongside text in message content blocks.

Built for the full execution path

The model's core directions pair tool calling and orchestration with trustworthy delivery. Execution status is meant to be transparent: the model attends to tool results and verifies outcomes rather than claiming completion without support.

Integration surface

Agnes 3.0 Flash is served through Chat Completions, Responses, and Messages APIs. A chat_template_kwargs extension field enables Thinking and other compatible features, and standard function-calling fields (tools, tool_choice) support agent workflows.

Benchmark Performance

Independent evaluations · Artificial Analysis

35.5%
Intelligence

Accuracy & Capability Details

GPQA - Graduate Science92.4%
Humanity's Last Exam38.5%
SciCode - Scientific Coding51.6%
Long Context Reasoning81.0%

Compare Models Side-by-Side

Evaluate specifications, pricing, and independent benchmark indices

Model Details
General Info
ProviderSapiens AIAnthropicMeta
Release DateSeptember 11, 2026September 1, 2026September 2, 2026
Knowledge Cutoff---
Context & Limits
Context Window512K-
1M
Best Context Window
Pricing (per 1M tokens)
Input Pricing
$0.05
Best Input Pricing
$10$1.25
Output Pricing
$0.15
Best Output Pricing
$50$4.25
Modalities
Inputs
textimage
textimagefile
textimagefilevideo
Outputs
text
text
text
Benchmarks (0-100)
Intelligence Index35.5
53.4
Best Intelligence Index
53.0
Coding Index-
81.6
Best Coding Index
76.3
Agentic Index-
58.0
Best Agentic Index
55.6
Agnes 3.0 Flash
Claude Fable 5.1
Muse Spark 1.3

GPQA Benchmark

Graduate-level reasoning and expert Q&A evaluation.

92%
Agnes 3.0 Flash
GPQA Benchmark
Score: 92%
Agnes 3.0 Flash
94%
Claude Fable 5.1
GPQA Benchmark
Score: 94%
Claude Fable 5.1
94%
Muse Spark 1.3
GPQA Benchmark
Score: 94%
Muse Spark 1.3

Humanity's Last Exam

Extremely difficult logical reasoning and knowledge.

39%
Agnes 3.0 Flash
Humanity's Last Exam
Score: 39%
Agnes 3.0 Flash
59%
Claude Fable 5.1
Humanity's Last Exam
Score: 59%
Claude Fable 5.1
49%
Muse Spark 1.3
Humanity's Last Exam
Score: 49%
Muse Spark 1.3

Long Context Reasoning

Logical reasoning over long context windows.

81%
Agnes 3.0 Flash
Long Context Reasoning
Score: 81%
Agnes 3.0 Flash
85%
Claude Fable 5.1
Long Context Reasoning
Score: 85%
Claude Fable 5.1
84%
Muse Spark 1.3
Long Context Reasoning
Score: 84%
Muse Spark 1.3

Independent evaluation data provided by Artificial Analysis. To view the latest benchmarks and full details, visit their official site.