NVIDIA's Vera Rubin platform moves agentic AI infrastructure into production across three layers at once: SpaceX and SpaceXAI are deploying the Vera CPU, Nebius is adopting the newly mass-produced Groq 3 LPX inference system, and CoreWeave is running Spectrum-X Multiplane networking that scales to 512,000 GPUs — turning the AI factory into one coordinated, full-stack system rather than a single-chip upgrade.
How does the Vera CPU become the core compute engine of the agentic AI factory?
SpaceX and SpaceXAI are the named early adopters driving Vera CPU into agentic AI production. SpaceX will deploy the Vera CPU at production scale for its next-generation agentic AI applicationsCITE:E5, and SpaceXAI has separately announced it will use the NVIDIA Vera CPU to power its next-generation agentic AICITE:E11.
On raw performance, NVIDIA compared its Vera Rubin NVL72 system against the prior-generation GB300 NVL72 using a DeepSeek V4 Pro agentic coding workload with context windows above 140,000 tokens. Depending on interaction speed, the gap in token throughput per megawatt between the two systems ranges from roughly 2x to 10x, reaching as high as 30x at the top endCITE:E1.
How does Groq 3 LPX reset AI inference performance records?
NVIDIA's Groq 3 LPX rack-level system has entered full mass productionCITE:E9. In agentic AI workloads, NVIDIA reports Groq 3 LPX has hit a record output speed of 2,900 tokens per secondCITE:E2.
On third-party benchmarks, the figures diverge slightly depending on the source. Inside reported that in Artificial Analysis's Gemma 4 31B (Reasoning) test at a 100,000-token context, Groq 3 LPX produced roughly 3,431 tokens per second — about 4x the second-place platformCITE:E3. NVIDIA's own blog cited a slightly different figure for the same use case, stating Groq 3 LPX reaches 3,400 tokens per second at a 100,000-token context, 4x faster than the closest alternative platformCITE:E10.
At the hardware level, a single rack-scale Groq 3 LPX deployment packs 256 LP30 accelerators, linked through direct chip-to-chip interconnect to function as one inference engine built for agentic workloadsCITE:E19.
How does Spectrum-X Ethernet support AI factories at million-GPU scale?
NVIDIA's Spectrum-X networking stack is what lets Vera Rubin racks scale into full AI factories, backed by a cluster of throughput, resilience, and switching numbers that NVIDIA has disclosed across both the Hot Chips 2026 presentation and its own blog.
| Metric | Value | Evidence |
|---|
| Scale-In VPC network performance gain | up to 18x | CITE:E6 |
| Scale-In storage network performance gain | 2x | CITE:E6 |
| Bandwidth retained with 1 of 8 network planes down | ~90% | CITE:E7CITE:E16 |
| Hardware failover recovery vs. software-based | up to 11x faster | CITE:E7CITE:E16 |
| AI factory Goodput / output gain | ~1.6x | CITE:E7CITE:E16 |
| Max GPU scale on Spectrum-X Multiplane | 512,000 GPUs | CITE:E8CITE:E15 |
| Spectrum-X networking performance vs. existing Ethernet | 1.6x | CITE:E14 |
| SN6000 switch ASIC throughput | 102.4 Tb/s | CITE:E17 |
| Max per-GPU bandwidth (ConnectX-9 SuperNIC) | 1,600 Gb/s | CITE:E17 |
| Spectrum-XGS multi-site NCCL acceleration | 1.9x | CITE:E18 |
The control plane for this architecture runs on the BlueField-4 ASTRA controllerCITE:E6. The Spectrum-X SN6000 switch family, built on a 102.4 Tb/s Spectrum-6 Ethernet ASIC paired with ConnectX-9 SuperNICs, is purpose-built for Vera Rubin NVL72 AI factoriesCITE:E17. Spectrum-X Multiplane's eight-plane topology means a single failed plane still leaves roughly 90% of total bandwidth intact, with hardware-based recovery running up to 11x faster than software-based multiplane load balancingCITE:E7CITE:E16, and it can flatten a two-tier network out to as many as 512,000 Rubin GPUs while cutting the number of switches requiredCITE:E8. For workloads that span multiple data centers, Spectrum-XGS Ethernet extends the same co-design across facilities so they can operate as a single AI super-factory, accelerating multi-site NCCL collective communication by 1.9xCITE:E18.
How are cloud providers accelerating adoption of NVIDIA's new AI factory platform?
CoreWeave and Nebius are the two cloud operators NVIDIA names as early production deployers of the new stack. CoreWeave is already running Spectrum-X Multiplane in production, using parallel switch banks to connect NVIDIA Vera Rubin racks into a high-bandwidth, flat, lossless AI networkCITE:E12.
On Groq 3 LPX cloud adoption, the two sources again differ slightly in framing. Inside reported that Nebius will become one of the first cloud providers to adopt Groq 3 LPXCITE:E4, while NVIDIA's own blog states Nebius is the first AI cloud to adopt NVIDIA Groq 3 LPXCITE:E13.
Taken together, the evidence shows three layers of the Vera Rubin stack moving into production in parallel rather than sequentially: Vera CPU compute at SpaceX and SpaceXAICITE:E5CITE:E11, Groq 3 LPX inference at NebiusCITE:E4CITE:E13, and Spectrum-X Multiplane networking at CoreWeaveCITE:E12. The throughput, recovery, and scale figures NVIDIA has disclosed — from the 512,000-GPU networking ceilingCITE:E8CITE:E15 to the 90%-bandwidth failover behaviorCITE:E7CITE:E16 — describe an architecture where compute, inference silicon, and networking are being deployed as one coordinated system rather than upgraded piece by piece.