Toolbit.aiToolbit.ai

Find, compare, and explore the best AI tools to match your specific tasks and use cases.

Explore

  • AI Search
  • Compare ToolsNew
  • Browse Categories
  • Trending Tools
  • Most Popular
  • New Additions

Resources

  • Updates HubNew
  • AI News
  • ModelsNew
  • Blog Articles
  • NewsletterNew

Company

  • Launch a Tool
  • Advertise with Us
  • Guest Post
  • Contact Us
© 2026 Toolbit.ai. All rights reserved.
Privacy PolicyTerms & ConditionsDisclaimer
There's An AI For That favicon
There's An AI For That•The front page of AI for everyone
Toolbit.ai
Toolbit.ai
UpdatesNew
Blog
Toolbit.ai
Toolbit.ai
Toolbit.ai
Toolbit.ai
UpdatesNew
Blog
Sign in
  1. Home
  2. Updates
  3. Models
  4. GLM 4.6V
Z AI
Released December 8, 2025

GLM 4.6V

GLM 4.6V by Z.AI (Zhipu AI) is a 106B vision-language MoE model with first-in-family native multimodal function calling, interleaved image-text generation.

Visit Z AIAnnouncement
Inputs
Image
Text
Video
Outputs
Text

Model Overview

Capabilities, design details, and architectural traits

GLM 4.6V – Native Multimodal Function Calling Vision-Language Model

GLM 4.6V is Z.AI's (Zhipu AI) 106B open-source vision-language model that introduces native multimodal function calling for the first time in the GLM-V family — bridging the gap between visual perception and executable action to provide a unified foundation for multimodal agents.

Native Multimodal Function Calling

GLM 4.6V's defining capability is its closed-loop perception-to-execution pipeline:

  • Multimodal Tool Input: Images, screenshots, and document pages are passed directly as tool call parameters — no prior text conversion required, preventing information loss.
  • Multimodal Tool Output: Visual results returned by tools (charts, rendered web screenshots, search-retrieved product images) are directly interpreted and incorporated into subsequent reasoning chains without text intermediaries.

Documented Capabilities

  • Interleaved Image-Text Generation: Takes multimodal context (documents, user inputs, tool-retrieved images) and synthesizes coherent, visually grounded mixed-media outputs. During generation, the model can autonomously call search and retrieval tools to curate additional visuals.
  • Multimodal Document Understanding: Processes up to 128K tokens of multi-document input as direct image pages - interpreting text, layout, charts, tables, and figures jointly without OCR preprocessing. Equivalent to ~150 complex document pages, 200 slide pages, or a 1-hour video in a single inference pass.
  • Frontend Replication & Visual Editing: Reconstructs pixel-accurate HTML/CSS/JS from UI screenshots by detecting layout, components, and color schemes. Supports natural-language-driven iterative visual edits - users can circle a region and issue a text instruction; the model locates and corrects the corresponding code.

Benchmark Performance

Independent evaluations · Artificial Analysis

16.9%
Intelligence

Accuracy & Capability Details

GPQA - Graduate Science71.9%
Humanity's Last Exam9.6%
SciCode - Scientific Coding30.4%
Instruction Following30.1%
Long Context Reasoning45.0%
τ²-Bench - Agentic Tasks31.6%
TerminalBench - System Control14.4%
Specs
Context window
131Ktokens
Input pricing
$0.30per 1M tokens
Output pricing
$0.90per 1M tokens

Prices in USD.

Compare Models Side-by-Side

Evaluate specifications, pricing, and independent benchmark indices

Model Details
General Info
ProviderZ AIAnthropicAnthropic
Release DateDecember 8, 2025July 24, 2026June 9, 2026
Knowledge Cutoff-May 2026-
Context & Limits
Context Window131K
1M
Best Context Window
1M
Best Context Window
Pricing (per 1M tokens)
Input Pricing
$0.30
Best Input Pricing
$5$10
Output Pricing
$0.90
Best Output Pricing
$25$50
Modalities
Inputs
imagetextvideo
textimage
textimagefile
Outputs
text
text
text
Benchmarks (0-100)
Intelligence Index10.9
63.1
Best Intelligence Index
62.1
Coding Index-
78.0
Best Coding Index
76.5
Agentic Index-
59.2
Best Agentic Index
56.6
GLM 4.6V
Claude Opus 5
Claude Fable 5

GPQA Benchmark

Graduate-level reasoning and expert Q&A evaluation.

57%
GLM 4.6V
GPQA Benchmark
Score: 57%
GLM 4.6V
93%
Claude Opus 5
GPQA Benchmark
Score: 93%
Claude Opus 5
93%
Claude Fable 5
GPQA Benchmark
Score: 93%
Claude Fable 5

Humanity's Last Exam

Extremely difficult logical reasoning and knowledge.

4%
GLM 4.6V
Humanity's Last Exam
Score: 4%
GLM 4.6V
55%
Claude Opus 5
Humanity's Last Exam
Score: 55%
Claude Opus 5
56%
Claude Fable 5
Humanity's Last Exam
Score: 56%
Claude Fable 5

Long Context Reasoning

Logical reasoning over long context windows.

14%
GLM 4.6V
Long Context Reasoning
Score: 14%
GLM 4.6V
76%
Claude Opus 5
Long Context Reasoning
Score: 76%
Claude Opus 5
77%
Claude Fable 5
Long Context Reasoning
Score: 77%
Claude Fable 5

Independent evaluation data provided by Artificial Analysis. To view the latest benchmarks and full details, visit their official site.

Back to all models