Key takeaways
- Vera Rubin NVL72 posted up to 3.7× higher Qwen3‑VL throughput than GB300 NVL72 in the MLPerf v6.1 preview.
- On DeepSeek‑R1 the same platform was up to 2.5× faster.
- GB300 NVL72 kept 99 % scaling efficiency when expanding from one to four 72‑GPU racks (288 GPUs total).
- Vera Rubin uses sixth-generation NVLink, while GB300 uses fifth-generation NVLink; NVIDIA rates Rubin's fabric at twice the per-GPU bandwidth and claims 10× higher packet rates and 3× lower latency than off-the-shelf Ethernet.
- Software refinements between v6.0 and v6.1 added up to 1.6× performance for GB300 on Qwen3‑VL.
- In a separate NVIDIA-reported SemiAnalysis AgentX preview—not MLPerf—Vera Rubin showed a 30× advantage over GB300.
1. Vera Rubin NVL72’s MLPerf v6.1 preview performance
NVIDIA announced on September 16 2026 that its first MLPerf Inference v6.1 submission for the Vera Rubin NVL72 platform delivered up to 3.7 times the throughput recorded for the GB300 NVL72 generation on the Qwen3‑VL benchmark.
For the DeepSeek‑R1 workload the same submission reported a 2.5 times advantage over GB300. MLCommons lists NVIDIA Rubin and Vera Rubin NVL72 in its Preview category because the systems are not yet generally available. MLCommons describes the v6.1 results as peer-reviewed; “preview” refers to platform availability, not an unverified benchmark run.
Software stacks used
- Qwen3‑VL – NVIDIA combined the open‑source vLLM inference engine with the Dynamo framework to drive the results.
- DeepSeek‑R1 – The workload was executed with TensorRT‑LLM.
The company attributes the gains to a full‑stack codesign that includes enhanced Tensor Cores, the Transformer Engine, and the NVFP4 precision format, which together shrink model memory footprints while preserving output quality.
2. GB300 NVL72 scaling efficiency
In the same MLPerf round NVIDIA demonstrated how the older GB300 NVL72 platform behaves when the GPU count is increased. A single rack containing 72 GPUs was compared against a four‑rack configuration totaling 288 GPUs on the DeepSeek‑R1 offline test, achieving 99 % scaling efficiency.
Throughput grew nearly in proportion to the hardware added. That result establishes near-linear scaling for this specific offline submission; it does not by itself isolate how much of the gain came from the interconnect, software, or other system components.
3. Two NVL72 generations, two NVLink fabrics
Vera Rubin NVL72 and GB300 NVL72 both organize 72 GPUs as a rack-scale domain, but they do not use the same NVLink generation. GB300 uses fifth-generation NVLink, rated at 1.8 TB/s of bidirectional bandwidth per GPU and 130 TB/s across the rack. Vera Rubin uses sixth-generation NVLink, with preliminary NVIDIA specifications of 3.6 TB/s per GPU and 260 TB/s per rack.
NVIDIA says sixth-generation NVLink provides 10× higher packet rates and 3× lower end-to-end latency than off-the-shelf Ethernet alternatives. Those are vendor comparisons for the Rubin fabric, not a measurement of NVLink's isolated contribution to the MLPerf gap. The 3.7× and 2.5× results compare complete hardware-software systems, so they cannot be attributed to the interconnect alone.
4. Software‑driven gains over the v6.0 submission
Even without new silicon, NVIDIA reported that software optimizations alone lifted GB300 performance by up to 1.6× on the Qwen3‑VL benchmark when moving from MLPerf v6.0 to v6.1. The improvements stem from:
- lower‑precision KV‑cache storage,
- additional kernel fusions,
- refined kernels, and
- disaggregated serving using vLLM and Dynamo.
Post‑submission tweaks (still unverified by MLCommons) have pushed token‑per‑second numbers higher for models such as GPT‑OSS‑120B and DLRMv3.
5. Agentic inference: MLPerf and AgentX are separate
MLPerf v6.1 introduced an Edge Agentic Inference test for multi-step workloads. Separately, NVIDIA reported that Vera Rubin NVL72 posted a 30× advantage over GB300 NVL72 in a SemiAnalysis AgentX preview. The AgentX figure is not an MLPerf v6.1 result and was not validated by MLCommons. NVIDIA also pointed to the upcoming MLPerf Endpoints benchmark as a future standardized measurement for API-based agentic workloads.
6. What the numbers could mean for infrastructure planners
The benchmark results give operators data points to consider when planning AI‑inference deployments. The 3.7× Qwen3‑VL boost indicates a potential for higher token throughput per rack, while the 99% scaling result shows that the submitted four-rack GB300 system maintained near-linear performance growth up to 288 GPUs. Software‑only gains of up to 1.6× demonstrate that performance can continue to improve without new hardware. Early agentic tests suggest that future AI services that rely on multi‑step reasoning may see larger relative benefits on the Vera Rubin platform, though actual economic impact will depend on specific workloads and deployment contexts.
Conclusion
The September 2026 MLPerf v6.1 results show Vera Rubin NVL72 reaching up to 3.7× the throughput of GB300 NVL72 on Qwen3-VL, but the benchmark does not isolate a single cause for that gap. GB300 scaled almost linearly to 288 GPUs with 99% efficiency, while software refinements lifted its Qwen3-VL performance by up to 1.6× over v6.0. NVIDIA's separate AgentX preview reported a 30× lead for Vera Rubin; planners should treat that as a vendor-reported early result, not an MLPerf-validated comparison.
Sources
This article was researched and fact-checked against the following sources:
- NVIDIA Vera Rubin NVL72: Up to 3.7x GB300 Performance (cloudnews.tech)
- NVIDIA Vera Rubin NVL72 Delivers Leading Performance in... (daily.dev)
- HPCwire - Since 1987 – Covering the Fastest Computers in the World and the People Who Run Them (hpcwire.com)
- NVIDIA Vera Rubin NVL72 Demonstrates Breakthrough Performance in MLPerf Inference v6.1 ~ More than Just SuccessFactors! (nageshpolu.com)
- NVIDIA Vera Rubin NVL72 Posts First MLPerf Inference Preview Results – Unite.AI (unite.ai)
- NVIDIA Vera Rubin NVL72 Delivers Leading Performance in MLPerf Inference v6.1 Debut (blogs.nvidia.com)
- MLCommons Sets Participation Record with New MLPerf Inference v6.1 Benchmark Results (mlcommons.org)
- NVIDIA NVLink and NVLink Switch (nvidia.com)