415.tech
AI & tech, from the frontlines of Silicon Valley
NVIDIA ships Groq 3 LPX, quoting 3,400 tokens/sec on 100K-token contexts

NVIDIA ships Groq 3 LPX, quoting 3,400 tokens/sec on 100K-token contexts

NVIDIA put Groq 3 LPX into full production as a rack-scale extension of Vera Rubin NVL72, quoting 3,400 output tokens per second on 100,000-token contexts running Gemma 4 31B — 4x the nearest platform in an Artificial Analysis benchmark, though that figure is NVIDIA's own framing of a single agentic workload. Nebius is the first AI cloud to adopt it and CoreWeave is running Spectrum-X Multiplane in production, so decode latency in long-context agent chains now has dedicated silicon behind it rather than general-purpose GPUs alone.

Source: blogs.nvidia.com

Post on XEmail

As AI agents reason, use tools and interact with other systems, they generate responses one token at a time, causing even tiny delays to multiply across complex chains of work.

NVIDIA

Why this matters

  • → Agentic AI decode latency now has dedicated silicon, eliminating the traditional GPU bottleneck for long-conte
  • → 4x faster token generation on 100K-context workloads shifts economics and responsiveness of multi-agent system
  • → Vera Rubin + Groq 3 LPX codesign shows infrastructure moving from general-purpose to purpose-built for reasoni
Decode latency gets silicon
Also in this edition