Not across MLPerf Inference as a whole. NVIDIA’s GB300 NVL72 Blackwell Ultra system delivered 45% higher throughput than its GB200 NVL72 comparison in the Offline DeepSeek-R1 test in MLPerf Inference v5.1, the specific result behind the “dominates” claim. But v5.1 is no longer the latest round: MLCommons had published v6.1 by October 5, 2026, and its analysis identifies NVIDIA Vera Rubin preview systems—not Blackwell Ultra—as having the largest per-accelerator gains on VLM and DeepSeek-R1. MLPerf is a collection of workload-specific results, not one overall score that establishes a universal winner.
What Blackwell Ultra’s MLPerf result actually shows
In its account of MLPerf Inference v5.1, NVIDIA said its GB300 NVL72 rack-scale system achieved 45% higher throughput than GB200 NVL72 on DeepSeek-R1 in the Offline scenario. This is a vendor-published account of a MLCommons-verified benchmark result. The figure describes one model, one scenario, one system comparison, and one benchmark round; it is not a claim that Blackwell Ultra is 45% faster for every inference task.
Offline measures throughput when a system can process a batch of requests without the same request-by-request latency constraints as an interactive service. That makes it useful for workloads where aggregate processing rate matters, but it should not be treated as a substitute for a Server or Interactive result. Nor does a rack-scale system comparison by itself establish how one accelerator performs relative to another.
How that claim fits the newer v6.1 results
MLCommons’ Inference v6.1 release, published in September 2026, is newer than the v5.1 Blackwell Ultra result. In its v6.1 analysis, MLCommons says the largest per-accelerator gains in VLM and DeepSeek-R1 came from NVIDIA Vera Rubin preview systems. For other tests, top results used hardware that also appeared in v6.0, with more gradual gains attributed to software-stack and algorithm improvements. The pattern is workload-specific: a new architecture can make a major difference on particular tests without leading every benchmark.
Free tools Windows power users keep installed
One-click scans. No signup required.
#1 Best Overall
- PLEASE NOTE: Exporting an NVIDIA RTX Pro 6000 GPU outside the US requires strict adherence to the U.S. Export Administration Regulations (EAR) and issuance of an export license from the Bureau of Industry and Security (BIS). Compliance and Know Your Customer (KYC) screening may be required as a condition of order acceptance. [NVIDIA Blackwell Streaming Multiprocessor] The new SM features increased processing throughput, and new neural shaders that integrate neural networks inside of programmable shaders | DLSS 4: Multi Frame Generation ensures ultra-smooth frame pacing for lifelike simulations.
- [Double-Flow-Through Design] The RTX PRO 6000 Blackwell features a double-flow-through cooling design, optimizing efficiency and airflow to sustain peak performance under 600W power loads. | [5th Gen Tensor Cores] Deliver up to 3X the performance of the previous generation and support for FP4 precision for faster AI model processing times with reduced memory usage, enabling local fine-tuning of LLMs and generative AI | [4th Gen Ray Tracing Cores] Double the ray-triangle intersection rate of the previous generation to create photoreal, physically accurate scenes and immersive 3D designs with RTX Mega Geometry, which enables up to 100X more ray-traced triangles.
- [PCIe Gen 5] Support for PCIe Gen 5 provides double the bandwidth of PCIe Gen 4, improving data-transfer speeds from CPU memory and unlocking faster performance for data-intensive tasks like AI, data science, and 3D modeling. | [GDDR7 Memory] With 96 GB of GPU memory and 1.8 TB ps bandwidth, it can tackle massive 3D and AI projects, fine-tune AI models locally, explore large-scale VR environments, and drive larger multi-app workflows.
- [DisplayPort 2.1] Achieve unparalleled visual clarity and performance, driving high resolution displays at up to 8K at 240 Hz and 16K at 60 Hz. Increased bandwidth enables seamless multi-monitor setups while HDR and higher color depth support ensures superior color accuracy for precision work, such as video editing, 3D design, and live broadcasting.
- [Universal MIG] Divide a single RTX PRO 6000 Blackwell into multiple isolated instances, each with dedicated resources, allowing for concurrent execution of multiple workloads, optimized GPU utilization, and secure isolation of different applications or users. [WARRANTY] 3 YR Manufacturer's Warranty. Bulk OEM Packaging. Retail Packaging is NOT included.
NVIDIA’s own v6.1 results summary reports up to 3.7 times the throughput for a Vera Rubin NVL72 preview submission compared with GB300 NVL72. Separately, NVIDIA reports 99% scaling efficiency for a four-system submission comprising 288 GB300 GPUs, as well as software advances of up to 1.6 times over v6.0. These are NVIDIA-attributed figures with different scopes: the Vera Rubin comparison is not a universal multiplier, and the scaling and software figures describe the stated GB300 submission and comparison.
| Result | What it compares | How to interpret it |
|---|---|---|
| 45% higher throughput | NVIDIA’s GB300 NVL72 versus GB200 NVL72 comparison on DeepSeek-R1 Offline in MLPerf Inference v5.1 (NVIDIA, 2025). | A result for that model, scenario, system comparison, and round—not an overall-suite lead. |
| Up to 3.7× throughput | NVIDIA’s Vera Rubin NVL72 preview submission versus GB300 NVL72 in its v6.1 summary (NVIDIA, 2026). | A vendor-reported comparison between those systems; it does not establish a multiplier for every benchmark or configuration. |
| Up to 5.7× best per-accelerator DeepSeek-R1 Server improvement | MLCommons’ comparison of the best result in v6.1 with the best result in v5.1 (MLCommons, 2026). | A across-round, per-accelerator comparison, not a Blackwell Ultra-only gain. |
| Up to 2.99× best per-accelerator VLM Server improvement | MLCommons’ comparison of the best result in v6.1 with the best result in v6.0 (MLCommons, 2026). | A across-round, per-accelerator comparison for VLM Server. |
The v6.1 figures should not be ranked against the 45% v5.1 result as if all were measurements of the same test. The models, scenarios, comparison baselines, and measurement bases differ.
Why MLPerf does not have one universal winner
MLCommons describes MLPerf Inference as an open-source, architecture-neutral and reproducible benchmark suite intended to provide technical information for tuning and procurement. In v6.1, the suite included 10 Datacenter and 6 Edge benchmarks, spanning scenarios such as Offline, Server, Interactive, SingleStream, and MultiStream. New tests included End-to-End RAG and Agentic Edge Inference; v6.1 also added an interactive VLM scenario and allowed speculative decoding in GPT-OSS-120B Interactive.
Rank #2
- Professional GPU with Blackwell Architecture
- Blackwell Architecture
- 24GB GDDR7 with PCIe 5.0 & Ray Tracing
- AI Workstation
Results are meaningful only within a compatible benchmark context. Model and scenario matter, as do the division, system configuration, and whether the reported rate is per accelerator or for the whole system. In the Closed division, a submission must use a model mathematically equivalent to the reference implementation, which holds the model fixed for a more direct comparison. The Open division allows different models or retraining, so its results answer a different kind of question.
MLCommons reported 30 participating organizations and 120 submitted systems across Datacenter and Edge, and Closed and Open, in v6.1. The participants included silicon vendors, system builders, cloud and neocloud providers, and inference-software specialists. Submitters choose which benchmarks to enter, so even a large leaderboard is not a single ranking of every system on every workload.
Keep system scale and availability in view
A system’s total throughput can rise with the number of accelerators, so whole-system rates and per-accelerator rates answer different questions. MLCommons’ v6.1 analysis says the largest submission used 512 accelerators: a Crusoe system that generated almost 5.8 million tokens per second on GPT-OSS-120B Offline. That is a large-system result on a different model and scenario from GB300’s v5.1 DeepSeek-R1 comparison; the raw token rates are not a like-for-like ranking.
Rank #3
- Form Factor: Plug-in Card
- Cooler Type: Active Cooler
- Maximum Power Consumption: 70W
- Length: 6.6
- Height: 2.7
MLPerf also classifies systems by availability. In MLCommons’ categories, Available systems can be purchased or rented in the cloud; Preview systems must be submittable as Available in the next round; and RDI systems are experimental, in development, or for internal use. A preview result can show where performance is heading, but does not establish that buyers can obtain the system under the same conditions today.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.How to use the benchmark for a buying decision
For an inference deployment, compare results that match the workload and operating goal rather than choosing a platform from a headline percentage. A practical comparison should align:
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
- Model and benchmark: use the same model and test, including the relevant quality target.
- Scenario: compare Offline with Offline, Server with Server, and Interactive with Interactive.
- Division: distinguish Closed results, which hold the model mathematically equivalent to the reference, from Open results that may use another model or retraining.
- Measurement basis: check whether throughput is per accelerator or for the entire system, and compare accelerator counts and system configurations accordingly.
- Availability: identify whether the system is Available, Preview, or RDI before treating the result as an option a buyer can deploy.
- Operating cost and power: account for whole-system configuration and power alongside throughput. MLCommons says its validated MLPerf Power figures refer to measured whole-system power for the accompanying benchmark.
That context is also why MLPerf is useful without being a complete procurement answer. As MLPerf Inference working-group co-chair Frank Han put it, “With performance data from the Inference v6.1 benchmark, customers can better understand the cost-benefit tradeoffs and make informed decisions on how to procure and deploy their AI systems.” The benchmark supplies comparable technical evidence when the test conditions match; buyers still need to relate those conditions to their own service, capacity, availability, and power requirements.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

