GLM-5.3-Flash
GLM-5.3-Flash by Z AI is the first natively multimodal GLM-5 model, pairing hybrid sparse-linear attention with a 1M context and visual coding loop.
Model Overview
Capabilities, design details, and architectural traits
GLM-5.3-Flash - The First Natively Multimodal GLM-5 Model
GLM-5.3-Flash is Z AI's first natively multimodal model in the GLM-5 series. Its defining idea is a hybrid architecture combining sparse and linear attention, the first of its kind in an open-source frontier model, keeping a 1M-token context window practical.
| Trait | Detail |
|---|---|
| Native multimodality | First model in the GLM-5 series with built-in vision: accepts text, image, video, and file input in one pass |
| Hybrid attention | Combines sparse and linear attention |
| Long context | 1M-token context window with up to 128K output tokens |
| Visual coding loop | Observes interfaces, rendered results, and interaction feedback to continuously test and improve its own work |
| Cross-environment coordination | Coordinates tasks across code, browsers, and GUIs, including real-world operation via BUA and CUA |
| Office deliverables | Autonomously breaks down goals, invokes tools, and produces finished PPTX, PDF, DOCX, and XLSX outputs |
| Always-on thinking | Thinking cannot be disabled; thinking.type only supports enabled |
Visual Intelligence Inside the Coding Loop
What separates GLM-5.3-Flash from text-only coding models is that vision is part of the workflow, not an add-on. For frontend, game, and 3D work, the model checks rendered output and iterates on what it sees. It can turn screenshots, multiple page images, website URLs, or screen recordings into complete applications, analyzing design systems, shared components, navigation structure, and animation logic before building, then comparing its own rendered pages against the references and refining differences.
The ox-alpha Stealth Debut
Z AI tested the model anonymously as ox-alpha on OpenCode and OpenRouter before it was officially revealed as GLM-5.3-Flash.
Benchmark Performance
Independent evaluations · Artificial Analysis
Accuracy & Capability Details
Compare Models Side-by-Side
Evaluate specifications, pricing, and independent benchmark indices
| Model Details | |||
|---|---|---|---|
| General Info | |||
| Provider | Z AI | Anthropic | Anthropic |
| Release Date | August 26, 2026 | September 22, 2026 | September 28, 2026 |
| Knowledge Cutoff | - | - | Jun 2026 |
| Context & Limits | |||
| Context Window | 1M | 1M | 1M |
| Pricing (per 1M tokens) | |||
| Input Pricing | $0.15 Best Input Pricing | $4 | $2 |
| Output Pricing | $0.50 Best Output Pricing | $20 | $10 |
| Modalities | |||
| Inputs | textimagevideo | textimagefile | textimagefile |
| Outputs | text | text | text |
| Benchmarks (0-100) | |||
| Intelligence Index | 41.8 | 57.6 Best Intelligence Index | 56.0 |
| Coding Index | 71.5 | - | - |
| Agentic Index | 50.9 | - | - |
Humanity's Last Exam
Extremely difficult logical reasoning and knowledge.
Long Context Reasoning
Logical reasoning over long context windows.
SciCode Benchmark
Scientific coding and mathematical modeling.
Independent evaluation data provided by Artificial Analysis. To view the latest benchmarks and full details, visit their official site.
Explore more from Z AI
Other models by Z AI
Top AI Models
Leading alternatives by intelligence score