Nvidia Corp. enters full production of Groq 3 LPX chips to accelerate agentic AI inference
Nvidia Corp. has officially entered full production of the Groq 3 LPX, a dedicated artificial intelligence inference accelerator designed to enhance the Vera Rubin data center platform. The new chip, developed through a $20 billion acquisition of Groq Inc., focuses on ultra-fast token generation to support highly responsive agentic AI workloads. By offloading the memory-bandwidth intensive decode phase to these accelerators, Nvidia Corp. can eliminate the tradeoff between throughput and response times, allowing AI agents to reason and execute complex tasks in real-time. In independent benchmarks, the Groq 3 LPX demonstrated a record-breaking 3,400 tokens per second when running the Gemma 4 31B model. This performance is four times faster than the nearest rival platform, according to Nvidia Corp. The chipmaker also revealed that neocloud provider Nebius Group N.V. will be the first customer to deploy the chips in its production inference platform. The administration announced that these chips will work in tandem with Nvidia's flagship GPUs to create a unified inference engine for enterprise-scale applications.
Sources
-
Nvidia's dedicated inference accelerator Groq 3 LPX enters full production to supercharge AI agents
SiliconANGLE
-
Nvidia says Groq racks will be online this year following $20 billion purchase
CNBC
-
Nvidia’s Groq Chip Will Shape AI Agent Usability
WSJ
-
NVIDIA Groq 3 LPX Now in Full Production With World-Class Speed for Agentic AI
NVIDIA Newsroom
-
What Nvidia's first Groq 3 LPU benchmarks tell us about its $20B gamble
The Register