NVIDIA's Nemotron 3 Nano 30B A3B: a hybrid Mamba-Transformer MoE model activating 3.2B parameters with configurable reasoning traces.
Capabilities, design details, and architectural traits
Nemotron 3 Nano 30B A3B utilizes an architectural design that merges sequence processing with selective associative recall. It leverages sparse activation to maintain broad knowledge while executing with the computational footprint of a much smaller network.
| Architectural Component | Implementation Details |
|---|---|
| Hybrid Layer Interleaving | Integrates 23 Mamba-2 and MoE layers for linear-time sequence processing with 6 Attention layers for precise recall. |
| Expert Routing Mechanism | Utilizes 128 standard experts plus 1 shared expert per MoE layer, routing each token to exactly 6 active experts. |
| Parameter Sparsity | Maintains a total base of 31.6B parameters while activating only 3.2B parameters per forward pass. |
| Configurable Reasoning Traces | Includes an enable_thinking toggle to bypass <think> block generation for low-latency inference workloads. |
The underlying foundation was pretrained using a specific Warmup-Stable-Decay learning rate schedule before undergoing multi-environment reinforcement learning to stabilize its reasoning outputs.
Independent evaluations · Artificial Analysis
Evaluate specifications, pricing, and independent benchmark indices
| Model Details | |||
|---|---|---|---|
| General Info | |||
| Provider | NVIDIA | Anthropic | Anthropic |
| Release Date | December 15, 2025 | July 24, 2026 | June 9, 2026 |
| Knowledge Cutoff | - | May 2026 | - |
| Context & Limits | |||
| Context Window | 256K | 1M Best Context Window | 1M Best Context Window |
| Pricing (per 1M tokens) | |||
| Input Pricing | $0.05 Best Input Pricing | $5 | $10 |
| Output Pricing | $0.20 Best Output Pricing | $25 | $50 |
| Modalities | |||
| Inputs | text | textimage | textimagefile |
| Outputs | text | text | text |
| Benchmarks (0-100) | |||
| Intelligence Index | 7.2 | 63.1 Best Intelligence Index | 62.1 |
| Coding Index | - | 78.0 Best Coding Index | 76.5 |
| Agentic Index | - | 59.2 Best Agentic Index | 56.6 |
Graduate-level reasoning and expert Q&A evaluation.
Extremely difficult logical reasoning and knowledge.
Logical reasoning over long context windows.
Independent evaluation data provided by Artificial Analysis. To view the latest benchmarks and full details, visit their official site.