GLM 4.6V by Z.AI (Zhipu AI) is a 106B vision-language MoE model with first-in-family native multimodal function calling, interleaved image-text generation.
Capabilities, design details, and architectural traits
GLM 4.6V is Z.AI's (Zhipu AI) 106B open-source vision-language model that introduces native multimodal function calling for the first time in the GLM-V family — bridging the gap between visual perception and executable action to provide a unified foundation for multimodal agents.
GLM 4.6V's defining capability is its closed-loop perception-to-execution pipeline:
Independent evaluations · Artificial Analysis
Evaluate specifications, pricing, and independent benchmark indices
| Model Details | |||
|---|---|---|---|
| General Info | |||
| Provider | Z AI | Anthropic | Anthropic |
| Release Date | December 8, 2025 | July 24, 2026 | June 9, 2026 |
| Knowledge Cutoff | - | May 2026 | - |
| Context & Limits | |||
| Context Window | 131K | 1M Best Context Window | 1M Best Context Window |
| Pricing (per 1M tokens) | |||
| Input Pricing | $0.30 Best Input Pricing | $5 | $10 |
| Output Pricing | $0.90 Best Output Pricing | $25 | $50 |
| Modalities | |||
| Inputs | imagetextvideo | textimage | textimagefile |
| Outputs | text | text | text |
| Benchmarks (0-100) | |||
| Intelligence Index | 10.9 | 63.1 Best Intelligence Index | 62.1 |
| Coding Index | - | 78.0 Best Coding Index | 76.5 |
| Agentic Index | - | 59.2 Best Agentic Index | 56.6 |
Graduate-level reasoning and expert Q&A evaluation.
Extremely difficult logical reasoning and knowledge.
Logical reasoning over long context windows.
Independent evaluation data provided by Artificial Analysis. To view the latest benchmarks and full details, visit their official site.