Nvidia's Groq inference chip hits production
Archive item — written before sources were shown.
Nvidia's $20B Groq deal shipped its first hardware: the Groq 3 LPX inference rack is in full production and heading to Nebius this year.
Nvidia announced August 24 that its Groq 3 LPX inference rack, the first hardware to ship from its $20 billion Groq licensing deal, is now in full production, with neocloud provider Nebius first in line to deploy it. Nvidia senior director Dion Harris told reporters the racks will be online later this year. The LPX rack pairs 32 compute trays of eight Groq LPUs each with Nvidia’s Vera Rubin GPU platform, and Nvidia says the combination delivers 3,400 output tokens per second on 100,000-token long-context workloads, roughly 4x faster than competing platforms running Gemma 4 31B.
Nvidia frames the chip as a decode-phase specialist that complements rather than replaces its GPUs: Groq’s SRAM-based design targets the low-latency token-generation step of serving a model, while Vera Rubin handles the rest. AWS separately said at GTC it will deploy Groq LPUs alongside more than a million Nvidia GPUs. The announcement also covered two related pieces of the same platform push: Spectrum-X Multiplane networking, which CoreWeave is deploying to scale to 512,000 GPUs without a third network tier, and NVLink Fusion, which lets chipmakers including Intel, MediaTek and Amazon’s Annapurna Labs connect custom XPUs to Nvidia’s rack architecture.
What it means for you
This is the first real product to come out of Nvidia’s pattern of licensing-and-hiring deals instead of acquisitions, and it answers a question that pattern left open: whether Groq’s technology would actually ship inside Nvidia’s stack or just sit on a balance sheet. If your agents spend a lot of time in the decode phase on long-context calls, a Groq-plus-GPU deployment from a cloud like Nebius or CoreWeave is now a concrete option to benchmark against a pure-GPU setup, not a future promise.
- 01With Groq 3 LPX in Full Production, NVIDIA Extends Vera Rubin Inference for Agentsblogs.nvidia.com · primary
- 02Nvidia says Groq racks will be online this year following $20 billion purchasecnbc.com · reporting
