SemiconductorsBRIEF

NVIDIA Vera Rubin: Hardware Gains and Open Nemotron 3 Ultra Benchmarks Target Post-Training Cost Efficiency for Agentic AI

N
NathanTechnology Editor · Technical Lead
Published · Updated
According to NVIDIA's blog, the Vera Rubin platform can train the largest models with one-fourth the GPUs used by the Blackwell generation, while Prime Intellect measured 30% greater throughput per CPU on Vera versus other x86 architectures. NVIDIA paired these hardware claims with an open-weight, 550-billion-parameter Nemotron 3 Ultra model that scored 71.7% on SWE-bench verified, and cited Perplexity, Together AI, and Prime Intellect as early post-training users.

How does Vera Rubin's hardware efficiency support NVIDIA's "intelligence per dollar" claim?

According to NVIDIA's blog, the NVIDIA Vera Rubin platform "trains the largest models with one-fourth the GPUs of the Blackwell generation" — meaning the same maximum model scale that previously required a full Blackwell-generation GPU fleet can, per NVIDIA's description, be trained with 25% of that GPU count on Vera Rubin.

On the CPU side, NVIDIA's blog cites third-party testing from Prime Intellect: when comparing realistic reinforcement-learning (RL) sandbox workloads against alternative x86 architectures, Prime Intellect found that the Vera CPU delivers, on average, 30% greater throughput per CPU. NVIDIA frames both figures — the one-fourth GPU ratio and the 30% per-CPU throughput gain — as the hardware basis for its "intelligence per dollar" argument for post-training workloads.

What are the open specs and benchmark results behind Nemotron 3 Ultra?

To back the efficiency claim with a concrete reference model, NVIDIA's blog points to NVIDIA Nemotron 3 Ultra, which it describes as "an open weight, 550-billion-parameter mixture-of-experts (MoE) model" that "offers verifiable benchmarks and a fully disclosed post-training recipe run on NeMo RL."

On performance, NVIDIA reports that Nemotron 3 Ultra "scored 71.7% on a standard real-world coding benchmark, SWE-bench verified, where it produced a working fix for roughly seven in 10 real software bugs from open source projects, each one checked against the project's own tests." NVIDIA presents the model's open weights and disclosed training recipe as the evidence base that lets outside parties verify the platform's post-training economics, rather than relying on unverifiable internal figures.

How is Perplexity using Vera Rubin-class infrastructure for post-training?

NVIDIA's blog describes Perplexity's RL post-training stack as running "asynchronously across hundreds of NVIDIA GPUs, with an RDMA-based weight transfer engine that syncs trillion-parameter models in under two seconds between training and inference compute nodes."

NVIDIA states that the resulting post-trained Qwen3 235B models are "then served on NVIDIA GB200 NVL72 systems." Together, the two data points describe a pipeline in which Perplexity's asynchronous RL training feeds directly into GB200 NVL72-based serving, with sub-two-second weight synchronization cited as the mechanism that keeps training and inference nodes aligned during that hand-off.

How large a post-training rollout capacity does NVIDIA project for Vera Rubin?

NVIDIA's blog offers an "illustrative 20 billion rollout tokens" figure for post-training on the Ultra-class model. NVIDIA states this number is "based on prior-generation Nemotron 3 Super's ~1.2 million rollouts at ~10,000 tokens each, scaled up for the larger Ultra model" — that is, the 20-billion-token estimate is derived arithmetically from the earlier Nemotron 3 Super's rollout volume (roughly 1.2 million rollouts, each around 10,000 tokens) and then scaled to account for Nemotron 3 Ultra's larger, 550-billion-parameter size described above.

NVIDIA labels the figure explicitly as illustrative rather than a measured production benchmark, distinguishing it from the directly measured GPU and CPU efficiency figures cited elsewhere in the same post.

Which companies has NVIDIA cited as adopting the Vera Rubin post-training platform?

Beyond Perplexity, NVIDIA's blog names Together AI, stating that it "provides post-training as a service, including supervised fine-tuning, RL and direct preference optimization," has "been running on NVIDIA's platform and optimized kernel libraries," and "is looking to harness the Vera Rubin platform next."

NVIDIA also credits Prime Intellect as the source of the CPU throughput testing described above — the same 30% per-CPU throughput advantage over alternative x86 architectures. Read together, NVIDIA's blog positions Prime Intellect as a party that has already benchmarked Vera hardware and Together AI as a post-training service provider planning to move onto Vera Rubin next, alongside Perplexity's already-running GB200 NVL72 deployment.

Key figures at a glance

MetricValueSource
Vera Rubin GPU requirement vs. Blackwell generationone-fourth the GPUsNVIDIA blog
Vera CPU throughput advantage (Prime Intellect testing)30% greater throughput per CPUNVIDIA blog
Nemotron 3 Ultra parameter count550 billion (MoE)NVIDIA blog
Nemotron 3 Ultra SWE-bench verified score71.7%NVIDIA blog
Perplexity weight-transfer sync time (trillion-parameter models)under two secondsNVIDIA blog
Illustrative Ultra-model rollout token estimate20 billion tokensNVIDIA blog
Nemotron 3 Super prior-generation rollout volume~1.2 million rollouts × ~10,000 tokens eachNVIDIA blog

What this means

All of the figures above come from a single NVIDIA blog post, so they should be read as NVIDIA's own account of its platform and partners rather than independently verified third-party reporting. Within that account, the pieces are internally consistent: the one-fourth GPU count and 30% per-CPU throughput gain describe the hardware layer; the open-weight, 550-billion-parameter Nemotron 3 Ultra with its 71.7% SWE-bench verified score serves as the reference workload NVIDIA uses to make those hardware claims checkable; and the 20-billion-token rollout estimate — itself derived by scaling Nemotron 3 Super's ~1.2 million rollouts — sizes the workload that the hardware is meant to run. Perplexity's sub-two-second weight sync and GB200 NVL72 serving of Qwen3 235B, together with Together AI's stated plan to move to Vera Rubin next and Prime Intellect's CPU benchmarking, are the concrete deployments and tests NVIDIA cites as evidence that this pipeline is already in use beyond NVIDIA's own lab.

📊 Evidence

FAQ

What model did NVIDIA use to demonstrate Vera Rubin's post-training economics?

NVIDIA cited Nemotron 3 Ultra, an open-weight, 550-billion-parameter mixture-of-experts model with a fully disclosed post-training recipe run on NeMo RL, which scored 71.7% on SWE-bench verified.

How fast does Perplexity sync models between training and inference nodes on its post-training stack?

According to NVIDIA's blog, Perplexity's RDMA-based weight transfer engine syncs trillion-parameter models in under two seconds between training and inference compute nodes, before serving the resulting Qwen3 235B models on GB200 NVL72 systems.

📎 Sources

  1. blogs.nvidia.com
N
NathanTechnology Editor · Technical Lead

Related

BRIEF

Intel Invests €5 Billion to Expand Manufacturing at Ireland's Leixlip Campus

According to Intel's newsroom announcement, the company is investing €5 billion ($5.7 billion) to expand its Leixlip campus in Ireland, producing Intel Xeon 6 and next-generation Xeon processors on the Intel 3 node. Intel says the capital programme began earlier this year and builds on more than €30 billion invested in Ireland since 1989.

Nathan ·
BRIEF

Meta Launches Muse Code, a Terminal-Based AI Coding Agent Built on Proprietary Muse Spark 1.2

According to TechCrunch and VentureBeat, Meta released Muse Code — a beta terminal-based AI coding agent for large repositories — alongside the Muse Spark 1.2 model on August 5, 2026. Muse Code installs with a single command, runs persistent background agents, and is entirely proprietary, putting Meta in direct competition with Anthropic's Claude Code and OpenAI's Codex.

Nathan ·
BRIEF

NVIDIA Opens Alpamayo 2 Super for Commercial Use in Autonomous Vehicles and Robotaxis

NVIDIA has moved Alpamayo 2 Super into commercial use, according to NVIDIA's official blog. The autonomous-driving reasoning model runs on Cosmos 3 Super Reasoner, holds three times the parameters of its 10-billion-parameter predecessors, and ships under the OpenMDW-1.1 license, which ITHome reports NVIDIA has extended across the entire Alpamayo model family for fine-tuning and commercial redistribution.

Nathan ·