NVIDIA says Vera Rubin NVL72 achieved up to 30× higher agentic inference throughput per megawatt and up to 35× lower cost per million tokens than GB300 NVL72 when running DeepSeek V4 Pro on the SemiAnalysis AgentX wor... The direct Vera Rubin comparison is with GB300 NVL72; NVIDIA separately says GB300 can deliver u...
Research answer

Create a landscape editorial hero image for this Studio Global article: What did Nvidia announce about its next-generation Vera Rubin NVL72 platform’s efficiency for agentic AI workloads—including its claimed inf. Article summary: NVIDIA’s headline claim is that Vera Rubin NVL72 can deliver up to 30× higher agentic-inference throughput per megawatt than GB300 NVL72 and reduce cost per token by up to 35×, based on its stated SemiAnalysis AgentX res. Topic tags: general, documentation, general web. Style: premium digital editorial illustration, source-backed research mood, clean composition, high detail, modern web publication hero. Use reference image context only for broad subject, composition, and topical grounding; do not copy the exact image. Avoid: logos, brand marks, copyrighted characters, real person likenesses, fake screenshots, UI text, readable text, watermarks, charts with fak
NVIDIA is positioning Vera Rubin NVL72 as an AI infrastructure platform built for workloads in which agents repeatedly generate tokens, call tools, execute code and create more context. In results published by NVIDIA, the system delivered up to 30× higher throughput per megawatt than GB300 NVL72 and reduced token cost by up to 35× on an AgentX test using DeepSeek V4 Pro.
Those are substantial claims, but they need to be read carefully: the supplied evidence does not provide a specific dollar cost per million tokens for Vera Rubin, nor does it establish an independent, apples-to-apples verification of the headline gains.
The headline comparison is:
The test is designed around agentic coding rather than a short, isolated question-and-answer exchange. AgentX replays production-style coding sessions that include context growth, tool calls and subagent spawning. 6
That distinction matters. An agent may generate an answer, inspect the result, call an external tool, execute code, create another task and continue reasoning with a much larger context. The infrastructure must therefore process more than the model’s initial response, increasing pressure on memory, networking, CPUs and power as well as on the accelerator itself.
The direct 30× claim compares Vera Rubin NVL72 with GB300 NVL72, not with Hopper. NVIDIA’s published AgentX-related material separately says GB300 NVL72 delivers up to 15× higher throughput per megawatt than an H200 Hopper system on DeepSeek V4 Pro.
That provides context for the generational progression, but the figures should not be multiplied into a single 450× claim. They are separate “up to” comparisons, and the available material does not supply one unified test table showing Vera Rubin, GB300 and Hopper under identical conditions.
The same limitation applies to cost. The evidence supports NVIDIA’s claim of a maximum 35× reduction versus GB300, but it does not provide a Vera Rubin dollar price per million tokens or a specific Vera-versus-Hopper cost figure.
NVL72 is a rack-scale system combining 72 Rubin GPUs and 36 Vera CPUs with ConnectX-9 SuperNICs and BlueField-4 DPUs. NVIDIA describes NVLink 6 as the scale-up fabric and Quantum-X800 InfiniBand and Spectrum-X Ethernet as the scale-out networking layers. 23
The design treats the rack as a tightly integrated computing system rather than a collection of independent accelerator servers. That is particularly relevant for agentic workloads, where performance depends on keeping accelerators supplied with data while coordinating CPU-side execution, networking and storage activity.
Vera is an 88-core CPU built around NVIDIA-designed Olympus cores. It uses Spatial Multithreading and LPDDR5X memory, with up to 1.2 TB/s of memory bandwidth. NVIDIA says Vera enables up to 1.8× faster task completion than x86 CPUs across workloads including agentic AI, reinforcement learning and data processing. 15
Its role is not simply to run a conventional host operating system. Vera is intended to handle the work surrounding model inference, including:
By accelerating these tasks, NVIDIA’s architecture aims to keep GPUs focused on inference while reducing bottlenecks in the rest of the agent loop. 159
BlueField-4 is NVIDIA’s dedicated infrastructure-processing component. The DPU combines a 64-core Grace CPU, LPDDR5X memory and ConnectX-9 networking, with connectivity of up to 800 Gb/s. 1
NVIDIA says BlueField-4 moves infrastructure functions—including control, security, data movement and orchestration—off the host CPUs and GPUs. In that model, the DPU acts as an infrastructure layer that helps provide isolated, secure and predictable operation across a large AI factory. 1
This is the logic behind NVIDIA’s “Scale-In” framing: efficiency is not limited to the GPU chip. It also depends on how effectively the rack handles communication, storage, scheduling and security without consuming accelerator capacity for those jobs.
DSX MaxLPS is software within NVIDIA’s broader DSX AI-factory approach, rather than a separate Vera Rubin GPU or CPU. NVIDIA describes it as a suite designed to maximize token performance per megawatt within a fixed power budget. Its stated approach combines liquid cooling and in-rack optimization so operators can run more GPUs near their energy-efficient operating points.
That makes DSX MaxLPS relevant to the 30× efficiency claim: the advertised metric is measured at the AI-factory level, where cooling, rack configuration and infrastructure software can affect tokens produced per unit of power. The available evidence does not establish how much of the reported improvement comes from the silicon, the software stack, the test configuration or the combination of all three.
NVIDIA and SpaceXAI said SpaceXAI plans to deploy Vera CPUs for its next generation of agentic-AI workloads and expand the infrastructure behind Grok using Vera Rubin. 45
The companies also described plans to extend an optimized Vera Rubin NVL72 system into orbit for a first-generation Starmind AI satellite. 4
These statements describe planned deployments. They are not evidence that an orbital NVL72-based system has already launched or entered service.
Vera Rubin NVL72’s significance is NVIDIA’s attempt to measure AI infrastructure against the way agents actually work: through long, iterative sessions that consume substantial context and involve tools, code and subagents. On that workload, NVIDIA reports up to 30× more throughput per megawatt and up to 35× lower token costs than GB300 NVL72.
For buyers and infrastructure teams, the key question is not whether the headline multiplier sounds large. It is whether the result holds under their own model, latency target, context length, concurrency, power budget and total-cost assumptions. The supplied evidence establishes NVIDIA’s claim and the architecture behind it, but not a final independent verdict.
Studio Global AI
This page includes a source-backed answer you can continue inside Studio Global.
NVIDIA says Vera Rubin NVL72 achieved up to 30× higher agentic inference throughput per megawatt and up to 35× lower cost per million tokens than GB300 NVL72 when running DeepSeek V4 Pro on the SemiAnalysis AgentX wor...
NVIDIA says Vera Rubin NVL72 achieved up to 30× higher agentic inference throughput per megawatt and up to 35× lower cost per million tokens than GB300 NVL72 when running DeepSeek V4 Pro on the SemiAnalysis AgentX wor... The direct Vera Rubin comparison is with GB300 NVL72; NVIDIA separately says GB300 can deliver up to 15× the throughput per megawatt of Hopper on the same DeepSeek V4 Pro class of workload.
The platform pairs 72 Rubin GPUs with 36 Vera CPUs, while BlueField 4 DPUs handle infrastructure work such as networking, security, orchestration and data movement.