According to NVIDIA's blog, the Vera Rubin platform can train the largest models with one-fourth the GPUs used by the Blackwell generation, while Prime Intellect measured 30% greater throughput per CPU on Vera versus other x86 architectures. NVIDIA paired these hardware claims with an open-weight, 550-billion-parameter Nemotron 3 Ultra model that scored 71.7% on SWE-bench verified, and cited Perplexity, Together AI, and Prime Intellect as early post-training users.
How does Vera Rubin's hardware efficiency support NVIDIA's "intelligence per dollar" claim?
According to NVIDIA's blog, the NVIDIA Vera Rubin platform "trains the largest models with one-fourth the GPUs of the Blackwell generation" — meaning the same maximum model scale that previously required a full Blackwell-generation GPU fleet can, per NVIDIA's description, be trained with 25% of that GPU count on Vera Rubin.
On the CPU side, NVIDIA's blog cites third-party testing from Prime Intellect: when comparing realistic reinforcement-learning (RL) sandbox workloads against alternative x86 architectures, Prime Intellect found that the Vera CPU delivers, on average, 30% greater throughput per CPU. NVIDIA frames both figures — the one-fourth GPU ratio and the 30% per-CPU throughput gain — as the hardware basis for its "intelligence per dollar" argument for post-training workloads.
What are the open specs and benchmark results behind Nemotron 3 Ultra?
To back the efficiency claim with a concrete reference model, NVIDIA's blog points to NVIDIA Nemotron 3 Ultra, which it describes as "an open weight, 550-billion-parameter mixture-of-experts (MoE) model" that "offers verifiable benchmarks and a fully disclosed post-training recipe run on NeMo RL."
On performance, NVIDIA reports that Nemotron 3 Ultra "scored 71.7% on a standard real-world coding benchmark, SWE-bench verified, where it produced a working fix for roughly seven in 10 real software bugs from open source projects, each one checked against the project's own tests." NVIDIA presents the model's open weights and disclosed training recipe as the evidence base that lets outside parties verify the platform's post-training economics, rather than relying on unverifiable internal figures.
How is Perplexity using Vera Rubin-class infrastructure for post-training?
NVIDIA's blog describes Perplexity's RL post-training stack as running "asynchronously across hundreds of NVIDIA GPUs, with an RDMA-based weight transfer engine that syncs trillion-parameter models in under two seconds between training and inference compute nodes."
NVIDIA states that the resulting post-trained Qwen3 235B models are "then served on NVIDIA GB200 NVL72 systems." Together, the two data points describe a pipeline in which Perplexity's asynchronous RL training feeds directly into GB200 NVL72-based serving, with sub-two-second weight synchronization cited as the mechanism that keeps training and inference nodes aligned during that hand-off.
How large a post-training rollout capacity does NVIDIA project for Vera Rubin?
NVIDIA's blog offers an "illustrative 20 billion rollout tokens" figure for post-training on the Ultra-class model. NVIDIA states this number is "based on prior-generation Nemotron 3 Super's ~1.2 million rollouts at ~10,000 tokens each, scaled up for the larger Ultra model" — that is, the 20-billion-token estimate is derived arithmetically from the earlier Nemotron 3 Super's rollout volume (roughly 1.2 million rollouts, each around 10,000 tokens) and then scaled to account for Nemotron 3 Ultra's larger, 550-billion-parameter size described above.
NVIDIA labels the figure explicitly as illustrative rather than a measured production benchmark, distinguishing it from the directly measured GPU and CPU efficiency figures cited elsewhere in the same post.
Which companies has NVIDIA cited as adopting the Vera Rubin post-training platform?
Beyond Perplexity, NVIDIA's blog names Together AI, stating that it "provides post-training as a service, including supervised fine-tuning, RL and direct preference optimization," has "been running on NVIDIA's platform and optimized kernel libraries," and "is looking to harness the Vera Rubin platform next."
NVIDIA also credits Prime Intellect as the source of the CPU throughput testing described above — the same 30% per-CPU throughput advantage over alternative x86 architectures. Read together, NVIDIA's blog positions Prime Intellect as a party that has already benchmarked Vera hardware and Together AI as a post-training service provider planning to move onto Vera Rubin next, alongside Perplexity's already-running GB200 NVL72 deployment.
Key figures at a glance
| Metric | Value | Source |
|---|
| Vera Rubin GPU requirement vs. Blackwell generation | one-fourth the GPUs | NVIDIA blog |
| Vera CPU throughput advantage (Prime Intellect testing) | 30% greater throughput per CPU | NVIDIA blog |
| Nemotron 3 Ultra parameter count | 550 billion (MoE) | NVIDIA blog |
| Nemotron 3 Ultra SWE-bench verified score | 71.7% | NVIDIA blog |
| Perplexity weight-transfer sync time (trillion-parameter models) | under two seconds | NVIDIA blog |
| Illustrative Ultra-model rollout token estimate | 20 billion tokens | NVIDIA blog |
| Nemotron 3 Super prior-generation rollout volume | ~1.2 million rollouts × ~10,000 tokens each | NVIDIA blog |
What this means
All of the figures above come from a single NVIDIA blog post, so they should be read as NVIDIA's own account of its platform and partners rather than independently verified third-party reporting. Within that account, the pieces are internally consistent: the one-fourth GPU count and 30% per-CPU throughput gain describe the hardware layer; the open-weight, 550-billion-parameter Nemotron 3 Ultra with its 71.7% SWE-bench verified score serves as the reference workload NVIDIA uses to make those hardware claims checkable; and the 20-billion-token rollout estimate — itself derived by scaling Nemotron 3 Super's ~1.2 million rollouts — sizes the workload that the hardware is meant to run. Perplexity's sub-two-second weight sync and GB200 NVL72 serving of Qwen3 235B, together with Together AI's stated plan to move to Vera Rubin next and Prime Intellect's CPU benchmarking, are the concrete deployments and tests NVIDIA cites as evidence that this pipeline is already in use beyond NVIDIA's own lab.