
Nvidia's open-weight Nemotron 3.5 Lightning hits 670 tokens/s, double Gemini 3.5 Flash-Lite
Nvidia's Nemotron 3.5 Lightning is a 31.6B-parameter hybrid Mamba-Transformer with 3.6B active parameters, scoring 24 on Artificial Analysis's Intelligence Index — level with gpt-oss-120b at a quarter the parameters. The real jump is speed and agentic work: nearly 670 tokens per second against Gemini 3.5 Flash-Lite's 386, and 824 Elo on GDPval-AA v2 versus gpt-oss-120b's 800. Qwen3.6 35B A3B (32) and Meta's Muse Glimmer (35) still score higher on raw intelligence, so the permissive OpenMDW-1.1 weights land as a cheap high-throughput engine for agent pipelines rather than a frontier model.
Source: the-decoder.com ↗
In pre-release tests using the final NVFP4 weights, the model hits nearly 670 tokens per second, the highest measured throughput among all compared models and almost twice as fast as Google's Gemini 3.5 Flash-Lite (386 tokens/s).
Why this matters
- → 670 tokens/sec beats Gemini Flash-Lite 2x, enables real-time agent pipelines
- → 31.6B params, 3.6B active — quarter the size of gpt-oss-120b at parity intelligence
- → Open weights + permissive license = free, deployable alternative to proprietary models