SemiconductorsBRIEF

NVIDIA Vera Rubin: 10x More Tokens Per Megawatt, One-Tenth the Cost Per Million Tokens, NVIDIA Says

N
NathanTechnology Editor · Technical Lead
Published · Updated
According to NVIDIA's July 21, 2026 blog post, Vera Rubin NVL72 delivers up to 10x more tokens per megawatt and one-tenth the cost per million tokens than GB200 NVL72. CoreWeave's DeepSeek-R1 benchmark and Google Cloud's new A5X instance each independently report the same 10x per-megawatt token multiplier on live hardware.

Why Vera Rubin Matters in the Agentic AI Era

According to NVIDIA's blog post, agentic AI systems — where models chain reasoning steps and call tools autonomously — can consume up to 15x more tokens than traditional AI applications, which NVIDIA frames as the reason efficient infrastructure has become a strategic priority for its partners (E11).

NVIDIA cited DeepInfra, an AI inference cloud platform, as processing nearly 5 trillion tokens a week, with about 30% of that volume driven by agentic systems (E20). Taken together, these two figures describe the demand backdrop NVIDIA uses to introduce Vera Rubin: a token-hungry workload category that did not exist at the same scale in prior AI infrastructure cycles.

How Vera Rubin NVL72 Achieves Its Cost-Per-Token and Performance-Per-Watt Claims

Three separate results, all reported in the same NVIDIA blog post dated July 21, 2026, converge on the same order of magnitude:

ComparisonMetricReported valueEvidence
CoreWeave DeepSeek-R1 benchmark vs. Grace Blackwell NVL72Tokens per megawatt10xE1
Vera Rubin NVL72 vs. GB200 NVL72Tokens per megawattup to 10xE12
Vera Rubin NVL72 vs. GB200 NVL72Cost per million tokensone-tenthE12
Google Cloud A5X vs. prior-generation instanceInference cost per tokenup to 10x lowerE18
Google Cloud A5X vs. prior-generation instanceToken throughput per megawattup to 10x higherE18

What Advantage Does the Vera CPU's Olympus Core Offer?

NVIDIA's blog post states that the Vera CPU's custom Olympus core delivers 2x single-threaded performance, 3x core-to-core bandwidth, and 40% lower memory latency versus competing chiplet designs (E3).

On the coordination side, NVIDIA cited DeepInfra benchmarks showing Vera CPU supports up to 1.6x more concurrent AI agents at the same quality of service and up to 2.2x faster orchestration than alternative CPUs, while improving infrastructure utilization and cost efficiency (E21).

Vera CPU metricReported advantageEvidence
Single-threaded performance2x vs. competing chiplet designsE3
Core-to-core bandwidth3xE3
Memory latency40% lowerE3
Concurrent AI agents (DeepInfra benchmark)up to 1.6x moreE21
Orchestration speed (DeepInfra benchmark)up to 2.2x fasterE21

How Do Sixth-Generation NVLink and the Spectrum Family Support Workloads at Rack and Multi-Site Scale?

Inside the rack, NVIDIA said sixth-generation NVLink scale-up delivers more than 2x throughput on complex workloads, 3x lower latency, and 10x higher packet rates than off-the-shelf Ethernet (E4). NVIDIA specified that the Vera Rubin NVL72's all-to-all NVLink 6 fabric reaches 260 TB/s, which it says removes bandwidth constraints so the rack functions as a single unified accelerator (E14).  

Between racks, Spectrum-X Ethernet combines 102.4T Spectrum-6 switch systems with 1.6T ConnectX-9 SuperNICs to reach 1.6x higher RDMA bandwidth than off-the-shelf Ethernet, per NVIDIA (E5). Across sites, NVIDIA said Spectrum-XGS Ethernet extends performance with 1.9x multi-site throughput (E7). CoreWeave's own deployment, built on the 102.4 Tb/s Spectrum-6 switch chip with a liquid-cooled design, delivers 1.64 Pb/s per rack100% more capacity than CoreWeave's previous-generation air-cooled switches, according to NVIDIA's blog post (E15).

InterconnectMetricReported valueEvidence
NVLink 6 (scale-up)Throughput vs. Ethernet2x+E4
NVLink 6 (scale-up)Latency vs. Ethernet3x lowerE4
NVLink 6 (scale-up)Packet rate vs. Ethernet10x higherE4
NVLink 6 (rack fabric)All-to-all bandwidth260 TB/sE14
Spectrum-X (Spectrum-6 + ConnectX-9)RDMA bandwidth vs. Ethernet1.6x higherE5
Spectrum-XGSMulti-site throughput1.9xE7
CoreWeave Spectrum-6 racksBandwidth per rack1.64 Pb/sE15
CoreWeave Spectrum-6 racksCapacity vs. prior air-cooled gen100% moreE15

What Do the Thermal and Power Optimizations Contribute?

NVIDIA introduced NVIDIA Photonics with co-packaged optics for scale-out, which it describes as the industry's first such switch in volume manufacturing, offering 5x lower power and 10x higher MTBI (mean time between interruptions) versus pluggable transceivers; NVIDIA named CoreWeave, Lambda, and OCI as among the first adopters (E6).

On cooling, NVIDIA said a 45-degree Celsius liquid-cooling inlet temperature design enables chiller-free, dry-cooler operation, and that for new AI factories this higher-temperature dry cooling combined with the closed-loop liquid-cooling system saves millions of gallons of water per megawatt annually (E9).

How Do Rack-Scale Codesign and the Global Supply Chain Support Faster Deployment?

NVIDIA said three generations of rack-scale codesign produced a Vera Rubin NVL72 system with no cables, fans, or hoses in the compute tray, cutting tray assembly time from hours to one minute (E8).

Behind that assembly speed, NVIDIA described a supply chain spanning 350+ factory sites in 30 countries, which it calls the largest, most mature rack-scale supply chain it has assembled to meet customer compute demand (E2).

How Are Microsoft, Mistral, CoreWeave, and Google Cloud Adopting Vera Rubin?

NVIDIA's post detailed several named deployments:

What This Means

Every efficiency figure in NVIDIA's post — CoreWeave's 10x tokens-per-megawatt result on DeepSeek-R1 (E1), NVIDIA's own 10x/one-tenth-cost comparison against GB200 NVL72 (E12), and Google Cloud's 10x cost and throughput figures for A5X (E18) — lands on the same order of magnitude, even though the three come from three different named parties measuring three different things (a live benchmark, a vendor-to-vendor spec comparison, and a cloud instance rollout). That consistency is notable, but all three figures were disclosed inside a single NVIDIA blog post; no independent, non-NVIDIA-affiliated benchmark appears in this evidence set. The demand rationale NVIDIA offers for this efficiency push — agentic workloads consuming up to 15x more tokens (E11), with DeepInfra already routing about 30% of its 5-trillion-token weekly volume through agentic systems (E20) — is likewise sourced entirely from the same post, via DeepInfra's own cited data. Readers should note that Microsoft/Mistral, CoreWeave, Google Cloud, and DeepInfra are named as adopters and benchmark partners, but the performance and cost figures attributed to them all appear in NVIDIA's own announcement rather than in separate releases from those companies.

📊 Evidence

FAQ

Which companies has NVIDIA named as early Vera Rubin adopters or benchmark partners?

According to NVIDIA's blog post, CoreWeave was the first AI cloud to bring up and validate Vera Rubin NVL72 and published the DeepSeek-R1 benchmark; Google Cloud launched its first A5X instance for London startup Ineffable Intelligence; Microsoft and Mistral signed a multibillion-dollar agreement covering thousands of Vera Rubin GPUs; and Lambda and OCI were named among the first adopters of NVIDIA Photonics.

What is the NVLink 6 fabric bandwidth inside a Vera Rubin NVL72 rack?

NVIDIA states the Vera Rubin NVL72's all-to-all NVLink 6 fabric reaches 260 TB/s, which it says lets the entire rack behave as a single unified accelerator.

How much more token volume does agentic AI generate compared with traditional AI applications, per NVIDIA?

NVIDIA's blog post says agentic systems can consume up to 15x more tokens than traditional AI applications, citing DeepInfra's data that about 30% of its nearly 5 trillion weekly tokens are already driven by agentic workloads.

📎 Sources

  1. blogs.nvidia.com
N
NathanTechnology Editor · Technical Lead

Related

BRIEF

At SIGGRAPH 2026, NVIDIA Announces Cosmos 3 Edge, DGX Station and MotionBricks for Agentic and Physical AI

According to NVIDIA's SIGGRAPH 2026 blog post, the company announced the 4-billion-parameter Cosmos 3 Edge model, DGX Station systems built on the GB300 Grace Blackwell Ultra Desktop Superchip with up to 20 petaflops of FP4 compute, a Synthetic Video Detector microservice reaching 92% accuracy on uncompressed video, and MCP-based AI agent integrations across five creative-software platforms.

Nathan ·
BRIEF

Intel's New Mexico Fab Pushes Packaging Reticle Limit to 8x Today, Targets 12x+ by 2028

According to Intel's newsroom announcement, Intel's Fab 9 advanced packaging site in Rio Rancho, New Mexico has scaled its 'reticle limit' to 8 times the current industry standard, with a target of over 12 times by 2028. Intel says the facility grew from 25 employees in 1980 to 2,700 employees and 500 suppliers today, using Foveros and EMIB/EMIB-T technologies to package AI chips.

Nathan ·