According to NVIDIA's July 21, 2026 blog post, Vera Rubin NVL72 delivers up to 10x more tokens per megawatt and one-tenth the cost per million tokens than GB200 NVL72. CoreWeave's DeepSeek-R1 benchmark and Google Cloud's new A5X instance each independently report the same 10x per-megawatt token multiplier on live hardware.
Why Vera Rubin Matters in the Agentic AI Era
According to NVIDIA's blog post, agentic AI systems — where models chain reasoning steps and call tools autonomously — can consume up to 15x more tokens than traditional AI applications, which NVIDIA frames as the reason efficient infrastructure has become a strategic priority for its partners (E11).
NVIDIA cited DeepInfra, an AI inference cloud platform, as processing nearly 5 trillion tokens a week, with about 30% of that volume driven by agentic systems (E20). Taken together, these two figures describe the demand backdrop NVIDIA uses to introduce Vera Rubin: a token-hungry workload category that did not exist at the same scale in prior AI infrastructure cycles.
How Vera Rubin NVL72 Achieves Its Cost-Per-Token and Performance-Per-Watt Claims
Three separate results, all reported in the same NVIDIA blog post dated July 21, 2026, converge on the same order of magnitude:
- CoreWeave's first benchmark on the DeepSeek-R1 model found Vera Rubin NVL72 produced 10x more tokens per megawatt than Grace Blackwell NVL72, which CoreWeave described, per NVIDIA's post, as "landing directly on the metric that matters most for power-constrained AI factories" (E1).
- NVIDIA said that compared with its own GB200 NVL72, Vera Rubin NVL72 delivers up to 10x more tokens per megawatt and one-tenth the cost per million tokens (E12).
- Google Cloud's A5X instances, announced at Google Cloud Next and built on Vera Rubin NVL72 rack-scale systems, deliver up to 10x lower inference cost per token and 10x higher token throughput per megawatt than the prior generation, according to NVIDIA (E18).
| Comparison | Metric | Reported value | Evidence |
|---|
| CoreWeave DeepSeek-R1 benchmark vs. Grace Blackwell NVL72 | Tokens per megawatt | 10x | E1 |
| Vera Rubin NVL72 vs. GB200 NVL72 | Tokens per megawatt | up to 10x | E12 |
| Vera Rubin NVL72 vs. GB200 NVL72 | Cost per million tokens | one-tenth | E12 |
| Google Cloud A5X vs. prior-generation instance | Inference cost per token | up to 10x lower | E18 |
| Google Cloud A5X vs. prior-generation instance | Token throughput per megawatt | up to 10x higher | E18 |
What Advantage Does the Vera CPU's Olympus Core Offer?
NVIDIA's blog post states that the Vera CPU's custom Olympus core delivers 2x single-threaded performance, 3x core-to-core bandwidth, and 40% lower memory latency versus competing chiplet designs (E3).
On the coordination side, NVIDIA cited DeepInfra benchmarks showing Vera CPU supports up to 1.6x more concurrent AI agents at the same quality of service and up to 2.2x faster orchestration than alternative CPUs, while improving infrastructure utilization and cost efficiency (E21).
| Vera CPU metric | Reported advantage | Evidence |
|---|
| Single-threaded performance | 2x vs. competing chiplet designs | E3 |
| Core-to-core bandwidth | 3x | E3 |
| Memory latency | 40% lower | E3 |
| Concurrent AI agents (DeepInfra benchmark) | up to 1.6x more | E21 |
| Orchestration speed (DeepInfra benchmark) | up to 2.2x faster | E21 |
How Do Sixth-Generation NVLink and the Spectrum Family Support Workloads at Rack and Multi-Site Scale?
Inside the rack, NVIDIA said sixth-generation NVLink scale-up delivers more than 2x throughput on complex workloads, 3x lower latency, and 10x higher packet rates than off-the-shelf Ethernet (E4). NVIDIA specified that the Vera Rubin NVL72's all-to-all NVLink 6 fabric reaches 260 TB/s, which it says removes bandwidth constraints so the rack functions as a single unified accelerator (E14).
Between racks, Spectrum-X Ethernet combines 102.4T Spectrum-6 switch systems with 1.6T ConnectX-9 SuperNICs to reach 1.6x higher RDMA bandwidth than off-the-shelf Ethernet, per NVIDIA (E5). Across sites, NVIDIA said Spectrum-XGS Ethernet extends performance with 1.9x multi-site throughput (E7). CoreWeave's own deployment, built on the 102.4 Tb/s Spectrum-6 switch chip with a liquid-cooled design, delivers 1.64 Pb/s per rack — 100% more capacity than CoreWeave's previous-generation air-cooled switches, according to NVIDIA's blog post (E15).
| Interconnect | Metric | Reported value | Evidence |
|---|
| NVLink 6 (scale-up) | Throughput vs. Ethernet | 2x+ | E4 |
| NVLink 6 (scale-up) | Latency vs. Ethernet | 3x lower | E4 |
| NVLink 6 (scale-up) | Packet rate vs. Ethernet | 10x higher | E4 |
| NVLink 6 (rack fabric) | All-to-all bandwidth | 260 TB/s | E14 |
| Spectrum-X (Spectrum-6 + ConnectX-9) | RDMA bandwidth vs. Ethernet | 1.6x higher | E5 |
| Spectrum-XGS | Multi-site throughput | 1.9x | E7 |
| CoreWeave Spectrum-6 racks | Bandwidth per rack | 1.64 Pb/s | E15 |
| CoreWeave Spectrum-6 racks | Capacity vs. prior air-cooled gen | 100% more | E15 |
What Do the Thermal and Power Optimizations Contribute?
NVIDIA introduced NVIDIA Photonics with co-packaged optics for scale-out, which it describes as the industry's first such switch in volume manufacturing, offering 5x lower power and 10x higher MTBI (mean time between interruptions) versus pluggable transceivers; NVIDIA named CoreWeave, Lambda, and OCI as among the first adopters (E6).
On cooling, NVIDIA said a 45-degree Celsius liquid-cooling inlet temperature design enables chiller-free, dry-cooler operation, and that for new AI factories this higher-temperature dry cooling combined with the closed-loop liquid-cooling system saves millions of gallons of water per megawatt annually (E9).
How Do Rack-Scale Codesign and the Global Supply Chain Support Faster Deployment?
NVIDIA said three generations of rack-scale codesign produced a Vera Rubin NVL72 system with no cables, fans, or hoses in the compute tray, cutting tray assembly time from hours to one minute (E8).
Behind that assembly speed, NVIDIA described a supply chain spanning 350+ factory sites in 30 countries, which it calls the largest, most mature rack-scale supply chain it has assembled to meet customer compute demand (E2).
How Are Microsoft, Mistral, CoreWeave, and Google Cloud Adopting Vera Rubin?
NVIDIA's post detailed several named deployments:
- Microsoft and Mistral expanded their partnership through a multibillion-dollar agreement to expand AI infrastructure in Europe; Mistral is adding capacity drawing on thousands of Vera Rubin GPUs for training, inference, and large-scale deployment (E10).
- CoreWeave became, per NVIDIA, the first AI cloud to bring up and validate Vera Rubin NVL72 after months of co-engineering work, and is now sharing the first measured performance numbers from live hardware — the DeepSeek-R1 benchmark cited above (E13).
- Google Cloud's first A5X instance is now running for London startup Ineffable Intelligence, powered by Vera Rubin NVL72, according to NVIDIA (E16). Ineffable Intelligence cofounder Lasse Espeholt said, as quoted in NVIDIA's post: "The next era of research requires the next era of hardware. We feel privileged to work with the teams at NVIDIA and Google Cloud, who were able to grant us early access to Vera Rubin. The support across both teams has been unmatched; we were up and running almost immediately and are already testing infra for our superlearners" (E17).
- NVIDIA said A5X uses ConnectX-9 SuperNICs combined with Google's next-generation Virgo networking, enabling clusters that scale to tens of thousands of Rubin GPUs within a single site and up to nearly a million GPUs across multisite configurations (E19).
What This Means
Every efficiency figure in NVIDIA's post — CoreWeave's 10x tokens-per-megawatt result on DeepSeek-R1 (E1), NVIDIA's own 10x/one-tenth-cost comparison against GB200 NVL72 (E12), and Google Cloud's 10x cost and throughput figures for A5X (E18) — lands on the same order of magnitude, even though the three come from three different named parties measuring three different things (a live benchmark, a vendor-to-vendor spec comparison, and a cloud instance rollout). That consistency is notable, but all three figures were disclosed inside a single NVIDIA blog post; no independent, non-NVIDIA-affiliated benchmark appears in this evidence set. The demand rationale NVIDIA offers for this efficiency push — agentic workloads consuming up to 15x more tokens (E11), with DeepInfra already routing about 30% of its 5-trillion-token weekly volume through agentic systems (E20) — is likewise sourced entirely from the same post, via DeepInfra's own cited data. Readers should note that Microsoft/Mistral, CoreWeave, Google Cloud, and DeepInfra are named as adopters and benchmark partners, but the performance and cost figures attributed to them all appear in NVIDIA's own announcement rather than in separate releases from those companies.