
NVIDIA puts Groq 3 LPX into production at a record 3,400 tokens per second
NVIDIA moved Groq 3 LPX, an extension of its Vera Rubin NVL72 platform, into full production; Artificial Analysis clocked a record 3,400 output tokens per second on Gemma 4 31B at a 100,000-token context. Nebius brings it to production first, in its Token Factory platform behind the same API — record generation rates on rented capacity, not owned racks. The 4x responsiveness figure over the nearest alternative is NVIDIA's own claim, not a third-party measurement.
Source: nvidianews.nvidia.com ↗
Inference is the growth engine of AI. NVIDIA Grace Blackwell and NVL72 revolutionized large language model inference with an unprecedented leap in performance and efficiency.
Jensen Huang, NVIDIA CEO
Why this matters
- → Agents can complete reasoning loops in minutes instead of hours
- → 4x faster than nearest alternative for latency-sensitive workloads
- → AI clouds now have production-ready infrastructure for agentic systems
Agents at extreme speed