Technology
Nvidia says Groq LPX racks are in full production and online later this year
The deployment follows Nvidia's $20 billion asset acquisition of the startup in December.
The short version
- Nvidia announced that its Groq 3 LPX racks are in full production following its $20 billion acquisition of Groq assets in December.
- The new hardware will be deployed alongside Vera CPUs and Rubin GPUs at neocloud provider Nebius, becoming operational later this year.
- The system targets low-latency inference for AI agents, allowing cloud providers to offer premium service tiers for time-sensitive tasks like coding.
Key facts
- Nvidia's Groq 3 LPX rack has entered full production following the company's $20 billion purchase of Groq assets in December.[CNBC]
- The Groq racks will be paired with Vera central processors and Rubin graphics processors at neocloud provider Nebius and are scheduled to go online later this year.[CNBC]
- Each LPX rack contains 256 Groq 3 chips, which are manufactured by Samsung.[CNBC]
- Nvidia cited an Artificial Analysis benchmark indicating the Groq 3 LPX rack can deliver 3,400 tokens per second.[CNBC]
- Competitors focusing on low-latency inference include AMD, which announced plans to integrate Cerebras chips into rack systems, and OpenAI, which uses Cerebras for its Ultrafast mode.[CNBC]
What remains uncertain
- The precise launch date for when the Groq racks will go live at Nebius later this year remains unspecified.[CNBC]