Nvidia confirmed Monday that its Groq 3 LPX racks are in full production. This step represents the first major commercialization of technology gained through the chipmaker's 20 billion dollar acquisition of Groq, which closed last December. The hardware will be installed at neocloud provider Nebius and is scheduled to come online later this year.
Solving for Latency
Nvidia senior director Dion Harris stated the racks will work alongside Vera central processors and Rubin graphics processors. Low latency stands as the primary goal for this architecture. AI agents need rapid response times to function well, particularly for complex tasks like software coding. Industry demand for such speed is high, and cloud providers can charge premium rates for the low-latency tokens these chips produce.
Harris noted this hardware provides a specific service for customers who require the most sensitive latency agreements. The Groq architecture uses 500 megabytes of SRAM directly on the chip die to cut down memory bottlenecks. While Nvidia relies on TSMC for its primary graphics processors, these specialized Groq chips are manufactured by Samsung.
Market Position and Competition
Each LPX rack contains 256 Groq 3 chips. Benchmarks from Artificial Analysis indicate these racks can deliver 3,400 tokens per second. Competition in this space remains active. Advanced Micro Devices recently announced it would connect its systems to chips from Cerebras, a company that also focuses on low-latency inference. OpenAI's recently announced Ultrafast mode promises 750 tokens per second using Cerebras hardware.
Nvidia management emphasized that these specialized chips serve a specific function rather than replacing the versatile GPU. GPUs remain the industry workhorse for both training and inference. The Groq technology targets the specific decode phase of model serving. This strategic split allows Nvidia to match the right processor to the right part of a workload.
Looking Ahead
Jensen Huang, Nvidia CEO, previously projected that cumulative sales for Blackwell and the new Vera Rubin systems will reach 1 trillion dollars by 2027. During the March product launch, Huang stated he would dedicate one quarter of data center space reserved for coding applications to Groq hardware. The rest of the space will feature Vera Rubin chips.
The deployment of these racks is part of a broader push to maintain dominance as AI infrastructure needs evolve. Nvidia holds its next earnings call this Wednesday, where investors expect more data on these hardware rollouts. The speed at which Nvidia moved to integrate Groq technology reflects the rapid pace of the current hardware cycle.

