GLM 4.5V by Z.AI (Zhipu AI) is a 106B vision-language MoE model with native thinking mode, multimodal grounding, and GUI/frontend replication.
Capabilities, design details, and architectural traits
GLM 4.5V is Z.AI's (Zhipu AI) 106B-parameter vision-language model (based on GLM-4.5-Air) that combines a sparse MoE language decoder with a pretrained vision encoder to support high-fidelity multimodal reasoning, document understanding, and GUI-agent workflows. It explicitly documents a thinking mode switch to trade deeper chain-of-thought style reasoning for latency and cost savings.
GLM 4.5V is available as an open model on Hugging Face (zai-org/GLM-4.5V) and is supported by inference toolchains (vLLM, Megatron Bridge) with FP8/BF16 runtimes; a smaller Flash variant is provided for low-latency or local deployment scenarios.
Independent evaluations · Artificial Analysis
Evaluate specifications, pricing, and independent benchmark indices
| Model Details | |||
|---|---|---|---|
| General Info | |||
| Provider | Z AI | Anthropic | Anthropic |
| Release Date | August 11, 2025 | July 24, 2026 | June 9, 2026 |
| Knowledge Cutoff | Dec 2024 | May 2026 | - |
| Context & Limits | |||
| Context Window | 66K | 1M Best Context Window | 1M Best Context Window |
| Pricing (per 1M tokens) | |||
| Input Pricing | $0.60 Best Input Pricing | $5 | $10 |
| Output Pricing | $1.80 Best Output Pricing | $25 | $50 |
| Modalities | |||
| Inputs | textimage | textimage | textimagefile |
| Outputs | text | text | text |
| Benchmarks (0-100) | |||
| Intelligence Index | 6.8 | 63.1 Best Intelligence Index | 62.1 |
| Coding Index | - | 78.0 Best Coding Index | 76.5 |
| Agentic Index | - | 59.2 Best Agentic Index | 56.6 |
Graduate-level reasoning and expert Q&A evaluation.
Extremely difficult logical reasoning and knowledge.
Logical reasoning over long context windows.
Independent evaluation data provided by Artificial Analysis. To view the latest benchmarks and full details, visit their official site.