NVIDIA pitches Vera Rubin on cost-per-token
NVIDIA argues its next-gen Vera Rubin platform maximizes intelligence-per-dollar, now backed by named cloud partners' own numbers.
NVIDIA published a blog post arguing its upcoming Vera Rubin platform delivers better intelligence-per-dollar than prior generations specifically for post-training and agentic-AI workloads, the compute-heavy fine-tuning and reinforcement-learning stages that follow initial model pretraining. The post is vendor benchmarking, not independent testing.
Update, July 21: NVIDIA published a follow-up with performance figures now attributed to specific cloud partners rather than NVIDIA’s own testing alone. CoreWeave reports 10x more throughput per megawatt than Grace Blackwell NVL72 on a DeepSeek-R1 workload; Google Cloud says its new A5X instances cut the cost of serving each token by roughly 90% while pushing token throughput per megawatt up tenfold over the prior generation; and versus NVIDIA’s own GB200 NVL72, NVIDIA cites up to 10x more tokens per megawatt at one-tenth the cost per million tokens. DeepInfra separately reported the platform’s new Vera CPU (88 custom Olympus cores) delivering 2.2x faster orchestration and support for 1.6x more concurrent agents at equivalent service quality. Microsoft Azure is also named as a validation partner, though without a comparable published figure yet.
Named partners publishing their own numbers is a meaningfully stronger form of evidence than a vendor’s internal benchmark alone, since CoreWeave, Google Cloud, and DeepInfra all have their own reputations riding on the figures they publish. It’s still not a substitute for benchmarking your own workload: “10x” claims cluster suspiciously tightly across three different partners and two different metrics, which is more consistent with a shared marketing brief than three independent measurements landing on the same number by coincidence. Read it as a strong signal Vera Rubin is a real efficiency step, not as a number you can plug directly into your own cost model.