NVIDIA confirmed on August 24, 2026 that its 256-chip Groq 3 LPX inference rack has entered full production, with Nebius named as the first cloud provider set to deploy it later this year alongside Vera CPU and Rubin GPU.
What Did NVIDIA Announce About the Groq 3 LPX Rack?
NVIDIA announced that its Groq 3 LPX AI inference rack has entered full productionCITE:E1. The company released the update at the Hot Chips 2026 conferenceCITE:E11, describing the rack as having reached a stage where performance hits industry-peak levels and calling it another milestone for the Vera Rubin platformCITE:E10. Investors expect Foxconn (鴻海, stock code 2317), identified as the largest assembler of related racks, to see shipments rise accordinglyCITE:E10.
Who Is the First Cloud Provider to Deploy Groq 3 LPX?
Nebius is the first AI cloud service provider to adopt NVIDIA's Groq 3 LPXCITE:E2. The rack will be paired with Vera CPU and Rubin GPU in Nebius's deployment, which is expected to go live later this yearCITE:E12.
What Is Inside the Groq 3 LPX Rack?
Each Groq 3 LPX rack packs 256 Groq 3 LPU chipsCITE:E4. NVIDIA licensed Groq's inference technology and recruited its core team last year, then unveiled the standalone Groq 3 LPU inference chip in March 2026 before rolling out the full rackCITE:E3. Each chip carries 500MB of embedded SRAM to reduce memory-driven compute bottlenecksCITE:E17; the SRAM is manufactured by Samsung Electronics, while the GPU portion of the platform is fabricated by TSMC (台積電, stock code 2330)CITE:E17.
How Fast Is Groq 3 LPX in Benchmark Testing?
In Artificial Analysis benchmark testing, Groq 3 LPX delivered response speeds up to four times faster than the most recent competing platforms on latency-sensitive workloadsCITE:E14. NVIDIA said the platform shows "world-class speed" in agentic coding and other latency-sensitive workloads, citing the same Artificial Analysis benchmarksCITE:E5. Running the Gemma 4 31B model with a 100,000-token context window, Groq 3 LPX processed 3,400 tokens per second, the highest figure recorded for that modelCITE:E16.
| Metric | Value | Source |
|---|
| Groq 3 LPU chips per LPX rack | 256 | CITE:E4 |
| Embedded SRAM per chip | 500MB | CITE:E17 |
| Latency-sensitive response speed vs. recent competing platforms | up to 4x | CITE:E14 |
| Token throughput (Gemma 4 31B, 100,000-token context) | 3,400 tokens/sec | CITE:E16 |
| Blackwell + Vera Rubin combined revenue, 2025–2027 | at least $1 trillion | CITE:E6 |
| Full-production announcement date | August 24, 2026 | CITE:E1 |
| NVIDIA FY2027 Q2 earnings date | August 26, 2026 | CITE:E9 |
How Does Groq 3 LPX Fit Into the Vera Rubin Platform?
Groq 3 LPX complements rather than replaces GPUs within NVIDIA's Vera Rubin architectureCITE:E15. Vera Rubin NVL72 handles the broader scope of training, inference, and context processing, while Groq 3 LPX is optimized specifically for interaction speed and token generationCITE:E18. NVIDIA CEO Jensen Huang (黃仁勳) said Vera Rubin extends this vision through AI factory configurations optimized for the agentic AI era, and that LPX's ultra-fast token generation pushes the performance frontier furtherCITE:E13.
What Other Adoption and Financial Milestones Surround This Launch?
SpaceXAI, led by Elon Musk, has adopted NVIDIA's Vera CPU to accelerate large-scale agentic AI computingCITE:E7. SpaceXAI plans to expand the AI infrastructure behind its Grok model through the Vera Rubin platform, and to extend an optimized Vera Rubin NVL72 system into space for the first-generation Starmind satellitesCITE:E8. Separately, Huang said in March 2026 at the GTC conference that combined revenue from the Blackwell and Vera Rubin product generations is projected to reach at least $1 trillion from 2025 through 2027CITE:E6. NVIDIA is scheduled to report its fiscal 2027 second-quarter earnings on August 26, 2026CITE:E9.
What This Means
The full-production announcement lands two days ahead of NVIDIA's August 26 earnings report, and it arrives alongside a separate SpaceXAI deal and Huang's $1 trillion Blackwell-plus-Vera-Rubin revenue projection covering 2025–2027CITE:E9CITE:E7CITE:E6. The supply chain named around the launch spans three points: Samsung producing the 500MB SRAM, TSMC fabricating the GPU, and Foxconn assembling the racksCITE:E17CITE:E10. At the same time, NVIDIA's own framing — that Groq 3 LPX supplements rather than replaces GPU-based inference on Vera Rubin NVL72 — positions the new rack as a companion product to the GPU business the revenue projection is built on, not a substitute for itCITE:E15CITE:E18.
Author's Take・EffectStory 編輯部
The specs disclosed here describe a clear division of labor rather than a GPU replacement: Vera Rubin NVL72 stays the training and general-inference workhorse, while Groq 3 LPX's 256-chip design and 500MB of embedded SRAM per chip exist to cut the latency tax on token-by-token, agentic interaction. The 4x response-speed edge and the 3,400-token-per-second record on Gemma 4 31B are the concrete evidence for that specialization, not a claim that LPX outperforms Rubin generally. What's worth watching next is whether Nebius's deployment, due later this year, and the August 26 earnings call show this second inference track scaling into the revenue base behind the $1 trillion Blackwell-to-Vera-Rubin projection — or whether it stays a narrower add-on for latency-sensitive workloads.