Do these 3 things before closing this tab:
1Scan for outdated or missing drivers - takes under a minute2Clear out junk files and repair common Windows errors3Fix the driver behind crashes, sound loss and screen glitchesEstimate an NVIDIA Vera Rubin NVL72 deployment from two inputs: a configuration-specific delivered quote and sustained, workload-specific benchmark results on the serving stack you intend to use. NVIDIA’s published peak throughput and vendor comparisons can help frame a capacity plan, but they do not tell you how many useful tokens per second your applications will serve or what those tokens will cost.
What does an NVL72 rack include, and what does its peak throughput tell you?
NVIDIA describes Vera Rubin NVL72 as a rack-scale system with 72 Rubin GPUs, 36 Vera CPUs, ConnectX-9 SuperNICs, BlueField-4 DPUs and NVLink 6. It is not 72 independent accelerator cards: the rack’s interconnect, networking and serving topology are part of the system being estimated. NVIDIA’s technical description specifies 18 compute trays, 9 NVLink switch trays, 3.6 TB/s bandwidth per GPU and 260 TB/s scale-up bandwidth per rack; these are NVIDIA-published specifications, not independent measurements. NVIDIA identifies Quantum-X800 InfiniBand and Spectrum-X Ethernet for scale-out. (NVIDIA product page and technical blog, 2026.)
| Published NVL72 figure | Format and workload label | How to use it |
|---|---|---|
| 3,600 PFLOPS | NVFP4 inference | Peak, format-specific arithmetic specification; not customer tokens per second. |
| 2,520 PFLOPS | NVFP4 training | Training specification, not an inference-capacity estimate. |
| 1,260 PFLOPS | FP8/FP6 training | Training specification. |
| 288 PFLOPS | FP16/BF16 | Published rack specification. |
| 144 PFLOPS | TF32 | Published rack specification. |
For a separate 100 MW AI-factory configuration using 40,000 Rubin GPUs with MaxLPS, NVIDIA lists 2 ZFLOPS of NVFP4 inference and 12 PB of HBM4. Those factory-scale figures should not be divided into a per-rack customer forecast without accounting for the configuration and deployment assumptions.
How much throughput can you use as a planning anchor?
Vendor comparisons are useful only when their model, context, software and test conditions match the decision at hand. NVIDIA’s product page reports one-tenth the cost per million tokens versus GB200 NVL72 and up to 10 times more tokens per megawatt for a Kimi-K2-Thinking comparison at 32K input and 8K output tokens. These are NVIDIA claims for that stated model and context, and NVIDIA says inference performance is subject to change. They are not a universal multiplier for other models or a substitute for measuring your costs.
Recommended Free Tools
#1 Best Overall
- NVIDIA Volta GV100 Architecture — 4,608 CUDA Cores, 640 1st-Gen Tensor Cores delivering 14 TFLOPS FP32 and 112 TFLOPS deep learning performance for AI training, inference, HPC, and scientific computing workloads
- 32GB HBM2 ECC Memory — 900 GB/s Bandwidth — High-bandwidth memory on a 4096-bit bus with ECC error correction provides the memory capacity and throughput required for the largest AI models, simulations, and datasets
- PCIe 3.0 x16 Interface — 250W TDP — Standard PCIe Gen3 connectivity with passive cooling designed for enterprise rack server deployment in HPE ProLiant, Dell PowerEdge, and Supermicro platforms with adequate chassis airflow
- NVLink — Scale to 96GB Unified Memory — Connect two V100 GPUs via NVLink at 300 GB/s bi-directional bandwidth to scale GPU memory from 32GB to 96GB for larger AI training and HPC workloads
- Multi-Precision Computing — Supports FP64 (7 TFLOPS), FP32 (14 TFLOPS), FP16 (112 TFLOPS) and INT8 precision modes for flexible deployment across training, inference, and scientific simulation workloads
NVIDIA’s separate training comparison—one-fourth as many GPUs in a projected MoE scenario involving a 10T model trained on 100T tokens in a fixed month—concerns training, not inference, so it should not enter an inference capacity or cost calculation.
In a September 16, 2026 post about MLPerf Inference v6.1 preview submissions, NVIDIA reported up to 3.7 times GB300 throughput for Qwen3-VL across offline, server and interactive scenarios using vLLM and NVIDIA Dynamo, and up to 2.5 times for DeepSeek-R1 using TensorRT-LLM. NVIDIA identified entries 6.1-0106 and 6.1-0074 and noted that optimization continued after submission. Treat these as NVIDIA-reported preview results for those workloads and stacks, not as an application-level multiplier for your deployment.
NVIDIA’s FY2026 Sustainability Report also presents modeled performance per megawatt using internal DLSim analytical projections for Vera Rubin NVL72 (NVFP4) with Groq 3 LPX (FP8) and GB200 NVL72 (NVFP4), including GPT-MoE 2T / 400K and a separate Kimi-K2-Thinking 32K/8K scenario. The report says modeled results may differ from measured silicon and other deployments.
How do you measure useful tokens per second?
Define the capacity boundary before benchmarking. A single NVL72 rack is not the same cost or performance boundary as a larger platform or POD that adds LPX, storage or context-memory systems, scale-out networking and adjacent racks. State which components are included in each benchmark and estimate.
- Fix the workload. Record the model and version, precision, prompt and output token distributions, context length, KV-cache assumptions, request mix and expected quality. Include tool calls or multistep agent requests if they are part of production traffic; NVIDIA CEO Jensen Huang described one agentic prompt as potentially launching a “thousand-step journey” of reasoning, retrieval, tool use and response generation in the company’s May 31, 2026 newsroom release. That is vendor framing, but it illustrates why a simple average prompt may not represent your request mix.
- Fix the serving setup. Record the serving framework and version, topology, batching and concurrency settings, and any other material software configuration. Benchmark the actual intended system scope rather than extrapolating from peak arithmetic figures.
- Set the service target. Specify latency targets, including time to first token and decode behavior where relevant, along with quality requirements and the service-level target. For long-context or interactive inference, measure prefill and decode separately as well as end-to-end results.
- Measure sustained output. Run representative traffic long enough to observe stable behavior at realistic concurrency. Record useful output tokens per second, latency distributions, achieved utilization and any rejected, timed-out or otherwise unusable output. Report these results together; throughput that misses the latency or quality target is not useful capacity for that service.
- Repeat across operating points. Measure the utilization levels you expect to sustain, not just a peak-load point. Retain the exact workload and configuration with each result so the cost model can be audited when software or traffic changes.
How much power does an NVL72 rack use?
NVIDIA’s product page does not publish a universal rack input-power figure. A 2026 Pegatron datasheet for its RA4803-72N3 Vera Rubin NVL72 implementation lists six 18.3 kW power supplies, “Max Q = 188kW” and “Max. TDP Support Max P = 228kW,” plus liquid cooling and 415V/480V input. Pegatron says specifications are subject to change. These are differently labeled, manufacturer- and configuration-specific values; they do not establish one sustained operating draw for every NVL72 deployment.
For a project estimate, ask the selected system builder and facility integrator to confirm sustained and maximum rack input power, the applicable electrical envelope, liquid-cooling capacity, distribution and redundancy requirements, and networking and adjacent-system loads. Do not treat chip TDP or a power-supply rating as measured full-rack consumption, or assume existing rack power and cooling are sufficient without an installation review.
How do you calculate total cost per million useful tokens?
Use a scenario with stated boundaries and assumptions, not a universal rack price. A practical annual model is:
Annual cost = annualized system and facility capital cost + annual energy and cooling cost + annual operations, service and support cost.
Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Repair Windows errors before they cause bigger problemsFix Now →Cost per million useful output tokens = annual cost ÷ (annual useful output tokens ÷ 1,000,000).
Use the sustained useful output measured under the service-level target to derive annual output for the hours the system is expected to operate. State operating hours, utilization and the assumed useful life. Finance teams may annualize delivered capital cost over a stated life using their accounting or financing method; include allocated facility fit-out or facility cost if it belongs in the chosen view. Add networking, storage, software, maintenance, staffing and support where applicable, and make clear what the quote includes.
For electricity, apply the local tariff to the expected energy use. If the facility uses a power-usage-effectiveness factor, state the value and apply it consistently to IT energy to account for facility overhead; avoid counting the same cooling or facility overhead again as a separate energy charge. Cooling infrastructure and its capital or service costs may still need their own explicit treatment in the project boundary.
Compare cost per million useful output tokens only when model quality, context, precision, latency, utilization and service-level assumptions are matched. A lower unit cost achieved by relaxing latency or quality, or by counting output that users cannot use, is not an equivalent result.
Rank #3
- Professional GPU with Blackwell Architecture in Compact Small Form Factor (SFF)
- Blackwell Architecture
- 24GB GDDR7 with PCIe 5.0 & Ray Tracing
- AI Workstation
What price and power figures are available publicly?
Public reporting offers estimates, not a standard purchase price. The following figures are dated, attributed examples and should not be treated as comparable quotes unless their configuration and included scope are known.
| Public figure | Attribution and date | What it establishes |
|---|---|---|
| About $7.8 million for a VR200 NVL72 rack | Tom’s Hardware, May 22, 2026, relaying a Morgan Stanley Research estimate | An analyst estimate reported secondhand, not NVIDIA’s selling price or a buyer quote. |
| “Up to $8.8 million” | Tom’s Hardware, March 24, 2026, based on secondary reporting | A separate reported upper estimate; configuration and scope are not established here as comparable to the other figure. |
| 188 kW and 228 kW maximum-related figures | Pegatron 2026 RA4803-72N3 datasheet | Manufacturer-specific labels for one implementation, not a universal sustained load. |
Use these public figures only to frame questions and rough scenarios. Obtain a dated quote with delivery, configuration, network and storage scope, software, support and any facility work itemized before presenting a buyer-specific estimate.
How should you compare rack ownership with cloud GPUs?
Public sources reviewed for this topic identify CoreWeave, Google Cloud, Microsoft Azure, Oracle Cloud Infrastructure and Nebius among providers deploying Vera Rubin systems, but they do not establish generally available customer prices or instance terms. Availability, geography, capacity reservation and service conditions must be confirmed directly for the intended buying date.
For a fair ownership-versus-rental comparison, benchmark or price the same model, quality, context length, precision, useful throughput, latency and utilization. Compare delivered system cost and facility power/cooling/network scope against the cloud provider’s current quote for the relevant region and service terms. Include the ownership estimate’s support, operations and facility costs, and compare it over the expected useful life. If you lack both a current provider quote and a comparable workload result, the evidence does not support a break-even claim.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
How should you handle uncertainty in the estimate?
Keep a low, base and high case for the inputs most likely to change the result: delivered rack quote, measured useful throughput, utilization, local electricity tariff and useful life. Make each input visible rather than hiding uncertainty inside a single cost-per-token figure. If the platform boundary, workload mix or facility assumptions differ between cases, state that explicitly so readers can see whether the result is driven by hardware economics, deployment scope or service expectations.
NVIDIA said on May 31, 2026 that Vera Rubin was ramping into full production and that production shipments were set to begin in fall 2026; its product and technical pages described production ramp and shipment plans in the second half of 2026. These company statements do not confirm a particular buyer’s delivery slot or regional availability, so obtain schedule and availability confirmation for the quoted configuration.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

