Celeris-1 by Celeris is a diffusion-based LLM generating 1,664 tokens/sec with 158ms median latency. It is 24x faster than GPT 5.
Capabilities, design details, and architectural traits
Celeris-1 is a diffusion LLM that replaces sequential autoregressive decoding with parallel refinement. Instead of predicting one token at a time, it starts with a rough version of the entire response and improves it over a few rapid passes, similar to how a blurry image comes into focus. This architecture is designed to preserve frontier-level reasoning while operating within real-time latency constraints.
| Trait | Detail |
|---|---|
| Diffusion architecture | Generates and refines the entire output sequence simultaneously across multiple passes, rather than token-by-token autoregressive decoding |
| 1,664 tokens/sec | Over 5x the output speed of Mercury 2 and 24x faster than GPT 5 |
| 158ms median latency | Responses land below the threshold at which humans perceive delay, enabling live conversational audio and real-time control loops |
| Second commercial diffusion LLM API | Follows Inception's Mercury as the second commercially available diffusion language model API |
Celeris-1 is engineered for agentic tool orchestration, classification, data extraction, and real-time voice pipelines. Its sub-perceptual latency compounds across multi-step agentic loops where dozens of internal reasoning calls for routing, validation, and tool selection each complete in milliseconds instead of seconds. The model is accessible through an OpenAI-compatible API, requiring only a base URL and key change to integrate into existing workflows.
Independent evaluations · Artificial Analysis
Evaluate specifications, pricing, and independent benchmark indices
| Model Details | |||
|---|---|---|---|
| General Info | |||
| Provider | Celeris | Anthropic | Anthropic |
| Release Date | July 24, 2026 | July 24, 2026 | June 9, 2026 |
| Knowledge Cutoff | - | May 2026 | - |
| Context & Limits | |||
| Context Window | - | 1M | 1M |
| Pricing (per 1M tokens) | |||
| Input Pricing | $0.20 Best Input Pricing | $5 | $10 |
| Output Pricing | $0.70 Best Output Pricing | $25 | $50 |
| Modalities | |||
| Inputs | text | textimage | textimagefile |
| Outputs | text | text | text |
| Benchmarks (0-100) | |||
| Intelligence Index | 12.4 | 63.1 Best Intelligence Index | 62.1 |
| Coding Index | 14.4 | 78.0 Best Coding Index | 76.5 |
| Agentic Index | 2.4 | 59.2 Best Agentic Index | 56.6 |
Graduate-level reasoning and expert Q&A evaluation.
Extremely difficult logical reasoning and knowledge.
Logical reasoning over long context windows.
Independent evaluation data provided by Artificial Analysis. To view the latest benchmarks and full details, visit their official site.