Nemotron 3.5 Lightning by NVIDIA is the fastest 30B MoE model for always-on agents, with hybrid Mamba-2 architecture, configurable reasoning, and 1M context.
Capabilities, design details, and architectural traits
NVIDIA Nemotron 3.5 Lightning is positioned as the fastest MoE model for completing specialized tasks for always-on agents. It supports context up to 1M tokens.
The model is built for customization. NVIDIA describes the BF16 reference weights as a starting point for post-training (SFT, RL, distillation), domain adaptation, and producing quantized variants. An NVFP4 release is provided separately for latency- and throughput-optimized inference.
| Trait | Detail |
|---|---|
| Architecture | Hybrid Mamba-2 + MoE + Attention layers (interleaved Mamba-2 and MoE with select Attention) |
| Reasoning Mode | Configurable on/off via chat template (enable_thinking=True/False) |
| Context Length | Up to 1M tokens |
| License | OpenMDW-1.1, ready for commercial use |
NVIDIA frames the BF16 release as reference weights for customization rather than optimized inference. The model is easily post-trained to achieve leading domain-specific accuracy and efficiency.
Independent evaluations · Artificial Analysis
Evaluate specifications, pricing, and independent benchmark indices
| Model Details | |||
|---|---|---|---|
| General Info | |||
| Provider | NVIDIA | Anthropic | Anthropic |
| Release Date | August 11, 2026 | July 24, 2026 | June 9, 2026 |
| Knowledge Cutoff | Sep 2025 | May 2026 | - |
| Context & Limits | |||
| Context Window | 1M | 1M | 1M |
| Pricing (per 1M tokens) | |||
| Input Pricing | Free Best Input Pricing | $5 | $10 |
| Output Pricing | Free Best Output Pricing | $25 | $50 |
| Modalities | |||
| Inputs | text | textimage | textimagefile |
| Outputs | text | text | text |
| Benchmarks (0-100) | |||
| Intelligence Index | 23.6 | 63.1 Best Intelligence Index | 62.1 |
| Coding Index | 26.8 | 78.0 Best Coding Index | 76.5 |
| Agentic Index | - | 59.2 Best Agentic Index | 56.6 |
Graduate-level reasoning and expert Q&A evaluation.
Extremely difficult logical reasoning and knowledge.
Logical reasoning over long context windows.
Independent evaluation data provided by Artificial Analysis. To view the latest benchmarks and full details, visit their official site.