415.tech
AI & tech, from the frontlines of Silicon Valley
NVIDIA puts Groq 3 LPX into production at a record 3,400 tokens per second

NVIDIA puts Groq 3 LPX into production at a record 3,400 tokens per second

NVIDIA moved Groq 3 LPX, an extension of its Vera Rubin NVL72 platform, into full production; Artificial Analysis clocked a record 3,400 output tokens per second on Gemma 4 31B at a 100,000-token context. Nebius brings it to production first, in its Token Factory platform behind the same API — record generation rates on rented capacity, not owned racks. The 4x responsiveness figure over the nearest alternative is NVIDIA's own claim, not a third-party measurement.

Source: nvidianews.nvidia.com

Post on XEmail

Inference is the growth engine of AI. NVIDIA Grace Blackwell and NVL72 revolutionized large language model inference with an unprecedented leap in performance and efficiency.

Jensen Huang, NVIDIA CEO

Why this matters

  • → Agents can complete reasoning loops in minutes instead of hours
  • → 4x faster than nearest alternative for latency-sensitive workloads
  • → AI clouds now have production-ready infrastructure for agentic systems
Agents at extreme speed
Also in this edition