NVIDIA Nemotron Nano 12B v2 VL is a vision-language model for multimodal document intelligence, multi-image reasoning, video understanding, visual Q&A.
Capabilities, design details, and architectural traits
Nemotron Nano 12B v2 VL is built for multimodal document intelligence. It is set up for image, video, and text inputs, with reasoning controlled by the system prompt for text and images.
| Distinct trait | Documented behavior |
|---|---|
| Multi-image reasoning | Supports reasoning over multiple images, including up to five input images. |
| Video understanding | Handles video inputs for tasks such as video understanding and video Q&A. |
| Prompt-controlled reasoning | Uses /think to enable reasoning and /no_think to disable it for text and images. |
| Document intelligence focus | Described for document tasks such as invoices, receipts, manuals, visual Q&A, and summarization. |
The model accepts image, video, and text inputs, and its documented use centers on multimodal document workflows. It also supports a 128K input plus output token context window.
The model returns text output and is presented as a commercial-use model. Its documented operation emphasizes document analysis and visual question answering rather than general chat.
Independent evaluations · Artificial Analysis
Evaluate specifications, pricing, and independent benchmark indices
| Model Details | |||
|---|---|---|---|
| General Info | |||
| Provider | NVIDIA | Anthropic | Anthropic |
| Release Date | October 28, 2025 | July 24, 2026 | June 9, 2026 |
| Knowledge Cutoff | - | May 2026 | - |
| Context & Limits | |||
| Context Window | 128K | 1M Best Context Window | 1M Best Context Window |
| Pricing (per 1M tokens) | |||
| Input Pricing | $0.20 Best Input Pricing | $5 | $10 |
| Output Pricing | $0.60 Best Output Pricing | $25 | $50 |
| Modalities | |||
| Inputs | imagetextvideo | textimage | textimagefile |
| Outputs | text | text | text |
| Benchmarks (0-100) | |||
| Intelligence Index | 4.2 | 63.1 Best Intelligence Index | 62.1 |
| Coding Index | - | 78.0 Best Coding Index | 76.5 |
| Agentic Index | - | 59.2 Best Agentic Index | 56.6 |
Graduate-level reasoning and expert Q&A evaluation.
Extremely difficult logical reasoning and knowledge.
Logical reasoning over long context windows.
Independent evaluation data provided by Artificial Analysis. To view the latest benchmarks and full details, visit their official site.