Agnes 3.0 Flash
Agnes 3.0 Flash by Sapiens AI is a text model for agentic coding and tool-driven tasks, with 512K context, multi-step tool orchestration, and verified
Model Overview
Capabilities, design details, and architectural traits
Agnes 3.0 Flash - Agentic coding and tool-driven execution model
Agnes 3.0 Flash is Sapiens AI's next-generation text model built for agentic coding and tool-driven tasks. Its defining focus is end-to-end execution quality: the full path from task understanding and planning, through tool use, to final delivery. The model is designed to make agent applications more reliable in complex, long-running work.
| Differentiator | What it means |
|---|---|
| Agentic coding focus | Tuned for Agnes Code and coding-agent workflows, strengthening execution from requirements understanding through final delivery. |
| Stable tool orchestration | Interprets tool definitions, selects appropriate tools, and coordinates multi-step calls while reducing ineffective calls, repetition, and abnormal loops. |
| Instruction and context adherence | Keeps objectives, constraints, and runtime context across long-running, multi-turn tasks, reducing task drift and missed requirements. |
| Trustworthy delivery | Emphasizes factual grounding and result verification to cut unsupported conclusions and incorrect completion claims. |
| Clean output | Reduced repetition, malformed text, and unnecessary exposure of internal reasoning, producing responses ready for delivery. |
| 512K context window | Supports long agent sessions with up to 65,536 output tokens. |
| Text and image-URL input | Accepts public image URLs alongside text in message content blocks. |
Built for the full execution path
The model's core directions pair tool calling and orchestration with trustworthy delivery. Execution status is meant to be transparent: the model attends to tool results and verifies outcomes rather than claiming completion without support.
Integration surface
Agnes 3.0 Flash is served through Chat Completions, Responses, and Messages APIs. A chat_template_kwargs extension field enables Thinking and other compatible features, and standard function-calling fields (tools, tool_choice) support agent workflows.
Benchmark Performance
Independent evaluations · Artificial Analysis
Accuracy & Capability Details
Compare Models Side-by-Side
Evaluate specifications, pricing, and independent benchmark indices
| Model Details | |||
|---|---|---|---|
| General Info | |||
| Provider | Sapiens AI | Anthropic | Meta |
| Release Date | September 11, 2026 | September 1, 2026 | September 2, 2026 |
| Knowledge Cutoff | - | - | - |
| Context & Limits | |||
| Context Window | 512K | - | 1M Best Context Window |
| Pricing (per 1M tokens) | |||
| Input Pricing | $0.05 Best Input Pricing | $10 | $1.25 |
| Output Pricing | $0.15 Best Output Pricing | $50 | $4.25 |
| Modalities | |||
| Inputs | textimage | textimagefile | textimagefilevideo |
| Outputs | text | text | text |
| Benchmarks (0-100) | |||
| Intelligence Index | 35.5 | 53.4 Best Intelligence Index | 53.0 |
| Coding Index | - | 81.6 Best Coding Index | 76.3 |
| Agentic Index | - | 58.0 Best Agentic Index | 55.6 |
GPQA Benchmark
Graduate-level reasoning and expert Q&A evaluation.
Humanity's Last Exam
Extremely difficult logical reasoning and knowledge.
Long Context Reasoning
Logical reasoning over long context windows.
Independent evaluation data provided by Artificial Analysis. To view the latest benchmarks and full details, visit their official site.