Z AI
Released August 26, 2026

GLM-5.3-Flash

GLM-5.3-Flash by Z AI is the first natively multimodal GLM-5 model, pairing hybrid sparse-linear attention with a 1M context and visual coding loop.

Inputs
Text
Image
Video
Outputs
Text

Model Overview

Capabilities, design details, and architectural traits

GLM-5.3-Flash - The First Natively Multimodal GLM-5 Model

GLM-5.3-Flash is Z AI's first natively multimodal model in the GLM-5 series. Its defining idea is a hybrid architecture combining sparse and linear attention, the first of its kind in an open-source frontier model, keeping a 1M-token context window practical.

TraitDetail
Native multimodalityFirst model in the GLM-5 series with built-in vision: accepts text, image, video, and file input in one pass
Hybrid attentionCombines sparse and linear attention
Long context1M-token context window with up to 128K output tokens
Visual coding loopObserves interfaces, rendered results, and interaction feedback to continuously test and improve its own work
Cross-environment coordinationCoordinates tasks across code, browsers, and GUIs, including real-world operation via BUA and CUA
Office deliverablesAutonomously breaks down goals, invokes tools, and produces finished PPTX, PDF, DOCX, and XLSX outputs
Always-on thinkingThinking cannot be disabled; thinking.type only supports enabled

Visual Intelligence Inside the Coding Loop

What separates GLM-5.3-Flash from text-only coding models is that vision is part of the workflow, not an add-on. For frontend, game, and 3D work, the model checks rendered output and iterates on what it sees. It can turn screenshots, multiple page images, website URLs, or screen recordings into complete applications, analyzing design systems, shared components, navigation structure, and animation logic before building, then comparing its own rendered pages against the references and refining differences.

The ox-alpha Stealth Debut

Z AI tested the model anonymously as ox-alpha on OpenCode and OpenRouter before it was officially revealed as GLM-5.3-Flash.

Benchmark Performance

Independent evaluations · Artificial Analysis

41.8%
Intelligence
71.5%
Coding Index
50.9%
Agentic Index

Accuracy & Capability Details

GPQA - Graduate Science91.2%
Humanity's Last Exam39.9%
SciCode - Scientific Coding51.6%
Long Context Reasoning80.0%

Compare Models Side-by-Side

Evaluate specifications, pricing, and independent benchmark indices

Model Details
General Info
ProviderZ AIAnthropicAnthropic
Release DateAugust 26, 2026September 22, 2026September 28, 2026
Knowledge Cutoff--Jun 2026
Context & Limits
Context Window1M1M1M
Pricing (per 1M tokens)
Input Pricing
$0.15
Best Input Pricing
$4$2
Output Pricing
$0.50
Best Output Pricing
$20$10
Modalities
Inputs
textimagevideo
textimagefile
textimagefile
Outputs
text
text
text
Benchmarks (0-100)
Intelligence Index41.8
57.6
Best Intelligence Index
56.0
Coding Index71.5--
Agentic Index50.9--
GLM-5.3-Flash
Claude Opus 5.5
Claude Sonnet 5.5

Humanity's Last Exam

Extremely difficult logical reasoning and knowledge.

40%
GLM-5.3-Flash
Humanity's Last Exam
Score: 40%
GLM-5.3-Flash
61%
Claude Opus 5.5
Humanity's Last Exam
Score: 61%
Claude Opus 5.5
55%
Claude Sonnet 5.5
Humanity's Last Exam
Score: 55%
Claude Sonnet 5.5

Long Context Reasoning

Logical reasoning over long context windows.

80%
GLM-5.3-Flash
Long Context Reasoning
Score: 80%
GLM-5.3-Flash
85%
Claude Opus 5.5
Long Context Reasoning
Score: 85%
Claude Opus 5.5
83%
Claude Sonnet 5.5
Long Context Reasoning
Score: 83%
Claude Sonnet 5.5

SciCode Benchmark

Scientific coding and mathematical modeling.

52%
GLM-5.3-Flash
SciCode Benchmark
Score: 52%
GLM-5.3-Flash
67%
Claude Opus 5.5
SciCode Benchmark
Score: 67%
Claude Opus 5.5
61%
Claude Sonnet 5.5
SciCode Benchmark
Score: 61%
Claude Sonnet 5.5

Independent evaluation data provided by Artificial Analysis. To view the latest benchmarks and full details, visit their official site.

Explore more from Z AI

Other models by Z AI

Top AI Models

Leading alternatives by intelligence score

View all