Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Estimate inference cost from the throughput your serving system sustains while meeting its latency target—not from a chip’s peak speed or hourly price alone. For cost per million output tokens, divide the system’s effective hourly cost by its measured output tokens per second, then multiply by 1,000,000 and divide by 3,600. The result is meaningful only when the cost and throughput cover the same capacity and workload.

What does an inference cost estimate measure?

A useful estimate answers a specific question: what does it cost to deliver a defined amount of inference for a particular model, under a particular service target? A chip name by itself cannot answer that. The result depends on the model and serving setup, how requests arrive, how much capacity is used, and what costs are included.

Choose the cost boundary first. An accelerator-only estimate includes the accelerator charge or its allocated ownership cost. A broader total cost of ownership (TCO) estimate may also include hosts, storage, networking, power, cooling, staffing, and availability-related costs. State which boundary you use, and apply it consistently to every option.

For the calculations below, the output-token rate is the number of generated tokens delivered by the serving system. If you instead report cost per million total tokens, define whether that includes both input and output tokens and use that same denominator throughout.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
Sale
HPE NVIDIA Tesla V100 32GB HBM2 PCIe 3.0 x16 Passive GPU Computational Accelerator for AI Machine Learning HPC Deep Learning 699-2G500-0216-400 (Renewed)
  • NVIDIA Volta GV100 Architecture — 4,608 CUDA Cores, 640 1st-Gen Tensor Cores delivering 14 TFLOPS FP32 and 112 TFLOPS deep learning performance for AI training, inference, HPC, and scientific computing workloads
  • 32GB HBM2 ECC Memory — 900 GB/s Bandwidth — High-bandwidth memory on a 4096-bit bus with ECC error correction provides the memory capacity and throughput required for the largest AI models, simulations, and datasets
  • PCIe 3.0 x16 Interface — 250W TDP — Standard PCIe Gen3 connectivity with passive cooling designed for enterprise rack server deployment in HPE ProLiant, Dell PowerEdge, and Supermicro platforms with adequate chassis airflow
  • NVLink — Scale to 96GB Unified Memory — Connect two V100 GPUs via NVLink at 300 GB/s bi-directional bandwidth to scale GPU memory from 32GB to 96GB for larger AI training and HPC workloads
  • Multi-Precision Computing — Supports FP64 (7 TFLOPS), FP32 (14 TFLOPS), FP16 (112 TFLOPS) and INT8 precision modes for flexible deployment across training, inference, and scientific simulation workloads

How do you measure AI inference cost-effectiveness?

1. Fix the workload

Keep the workload constant when comparing accelerators. Record the model and version, input-to-output token mix, context length, request arrival pattern, concurrency, precision or quantization, serving software, and deployment mode. A change in any of these can change both throughput and latency.

2. Set the latency target

Decide what responsiveness the service must deliver before measuring throughput. Set latency limits and the percentile that matters to users; for interactive generation, track time to first token and time per output token where relevant. Google Cloud’s AI accelerator performance and benchmarking guidance recommends increasing concurrent requests and stopping when P99 latency violates the service-level agreement (SLA). Record sustained throughput at the last operating point that meets the target, rather than using an unconstrained saturation result.

3. Measure delivered throughput at that operating point

Measure output tokens per second for the complete serving system, and record the per-accelerator rate if it helps compare systems of different sizes. Include the concurrency and latency results alongside throughput. Peak throughput that misses the service target is not usable capacity for that service.

Rank #2
MX3 M.2 AI Accelerator
  • High-Performance AI Processing: The MX3 is designed to handle the most demanding AI computer vision workloads, delivering exceptional performance and efficiency.
  • Flexible Integration: The MX3 can be easily integrated into your existing systems via its M.2 M-key form factor and support for Linux operating systems.
  • Energy Efficient: The MX3 is designed to provide high performance while minimizing power consumption.
  • Comprehensive Software Development Kit (SDK): The MX3 is supported by a comprehensive SDK that simplifies development and deployment.
  • Hardware compatability: The MX3 is compatible with the PCI-SIG M.2 M-key 2280 Specification. It can be used with the Raspberry Pi 5 with a M-key 2280 HAT.

Google Cloud’s benchmarking guidance recommends fixed-model comparisons, concurrency sweeps, latency targets, and sustained throughput per chip. It also discusses training; training examples should not be treated as inference estimates.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How do you calculate cost per million output tokens?

Let C be the effective cost in dollars per hour and T the measured, sustained output rate in tokens per second at the chosen service target:

Cost per million output tokens = C × 1,000,000 ÷ (T × 3,600)

Rank #3
waveshare Hailo-8 M.2 AI Accelerator Module, Compatible with Raspberry Pi 5, Supports Linux/Windows Systems, Based On The 26TOPS Hailo-8 AI Processor, Module Only
  • ✅Powered by 26 Tera-Operations Per Second (TOPS) Hailo-8 AI Processor. 2.5W typical power consumption
  • ✅Scalable, enabling simultaneous processing of multi-streams & multi-models
  • ✅Enabling real-time, low latency and high-efficiency AI inferencing on the edge devices
  • ✅Supports TensorFlow, TensorFlow Lite, ONNX, Keras, Pytorch frameworks
  • ✅Supports Linux and Windows. Supports the temperature range of -40°C to 85°C

The 3,600 converts an hour into seconds; multiplying the hourly rate by 1,000,000 scales the result to a million output tokens. This is a unit conversion, not a published benchmark. Use aggregate system throughput with the system’s hourly cost. If calculating per accelerator, use the per-accelerator throughput and matching per-accelerator cost. Mixing a whole-system price with a per-chip rate—or the reverse—produces a misleading figure.

Choose the right hourly cost

For rented infrastructure, use the actual product- and region-specific charge and the billed unit that corresponds to the capacity you measured. Account for commitment or pricing terms if they apply. For owned equipment, estimate an effective hourly cost from purchase or lease cost spread over useful life, then include the ongoing costs inside your stated TCO boundary. NVIDIA’s 35x Lower Token Cost with Blackwell describes the cloud numerator as the provider’s hourly rate and the owned-infrastructure numerator as an effective hourly cost derived from amortization; it also cautions that compute price or FLOPs per dollar alone gives an incomplete view of inference TCO.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How should you account for utilization and changing demand?

A saturated benchmark describes a measured operating point, not necessarily the cost of serving real traffic. If you pay for capacity that sits idle, fewer tokens are delivered per paid hour and effective cost per token rises. Conversely, a system that meets a latency target at a higher sustained request rate may spread its hourly cost across more delivered tokens.

Rank #4

Run a load sweep at low, typical, and peak expected request rates. At each level, record the delivered output tokens, latency percentiles, and paid capacity. This reveals whether a single high-load measurement reflects actual use, and whether demand peaks require extra capacity that is underused at other times.

A June 2026 arXiv preprint by Chitral Patil reports costs from $0.21 to $15.25 per million output tokens across tested conditions on identical H100 hardware. That range is specific to the paper’s model, serving setup, and load conditions; it is not a general H100 cost estimate or a multiplier to apply to another deployment.

How do you compare cloud and owned accelerator options?

Compare like with like: use the same model, token mix, serving software, precision, latency target, and demand profile where possible. Record the cost boundary and the units of both the price and capacity. For each option, retain the measured operating point rather than comparing one system’s latency-constrained throughput with another’s peak result.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Best Value
ASRock Radeon AI PRO R9700 Creator 32GB Professional Graphics Card, 2920 MHz Boost Clock, GDDR6, AMD RDNA 4, AI-Accelerators, DisplayPort 2.1a, PCIe 5.0, Blower Cooler
  • Professional AI & Creator Workstation: AMD Radeon AI PRO R9700 GPU with 32GB GDDR6 is engineered for AI development, professional content creation, and compute-intensive workloads.
  • Massive 32GB Memory Capacity: 32GB of GDDR6 memory on a 256-bit bus provides ample bandwidth for large AI models, 8K video editing, and complex 3D rendering.
  • Advanced RDNA 4 with AI Accelerators: 64 Compute Units with 3rd Gen Ray Tracing and dedicated 2nd Gen AI Accelerators for groundbreaking AI performance and visual computing.
  • Professional Blower Cooling: Efficient single blower design exhausts heat directly out of the chassis, ideal for multi-GPU workstation and server configurations.
  • Enterprise-Grade Thermal Solution: Vapor chamber heatsink with industrial Honeywell PTM7950 thermal interface material ensures reliable cooling under sustained professional loads.
Comparison item What to record
System and deployment Accelerator and system configuration; rented or owned; region and pricing terms for rented capacity.
Workload and serving Model and version, input/output mix, context length, precision or quantization, serving software, and deployment mode.
Service point Sustained output tokens per second, concurrency, time to first token or time per output token as relevant, and latency percentiles.
Cost and utilization Hourly charge or effective owned hourly cost; billed unit; cost boundary; low, typical, and peak load results.
Derived result Cost per million output tokens calculated from matching cost and throughput units.

Cloud GPUs, cloud TPUs, and owned infrastructure are all possible deployment choices, but the figures below do not establish a universal winner. Google Cloud’s benchmarking material also recommends testing representative model types: include a dense model and, if relevant to the deployment, a sparse or mixture-of-experts (MoE) model or a reasoning model.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What do published accelerator prices and benchmarks show?

Cloud list prices and vendor benchmark figures can help illustrate the inputs to an estimate, but they are not substitutes for workload-specific measurement.

Published figure What it represents Important qualification
$12.00 per chip-hour for Ironwood Google Cloud on-demand price in us-central1 (Iowa), on its pricing page accessed in 2026. Product-, region-, and date-specific live pricing; verify the current price and billed unit before estimating.
$2.70 per chip-hour for Trillium Google Cloud on-demand price in us-east1 (South Carolina), on its pricing page accessed in 2026. Product-, region-, and date-specific live pricing; verify the current price and billed unit before estimating.
$4.20 per chip-hour for TPU v5p Google Cloud on-demand price in us-east5 (Columbus), on its pricing page accessed in 2026. Product-, region-, and date-specific live pricing; verify the current price and billed unit before estimating.
$4.20 per million tokens for an H200 example; $0.12 per million tokens for a GB300 NVL72 example A 2026 comparison published by NVIDIA. These are comparison-specific vendor figures, not general prices for those accelerators or independently established universal rankings.
$0.123 per million tokens for GB300 NVL72 at 116 tokens per second per user NVIDIA’s figure citing SemiAnalysis InferenceX, as of April 2026. A benchmark-specific claim; retain the stated per-user throughput condition when citing it.

Google Cloud says TPU charges accrue while a node is in READY state and lists prices per chip-hour. A TPU VM can contain multiple chips, while console billing may appear in VM-hours. Confirm that the price and usage quantity use matching units; a VM-hour is not interchangeable with a chip-hour unless the VM’s chip count is accounted for. Recheck regional pricing because the listed rates can change.

NVIDIA’s cost and benchmarking pages cite selected systems and benchmarks, including SemiAnalysis InferenceX and MLPerf Inference. A result from one comparison should be carried forward only with its named workload, system configuration, software, latency or interactivity condition, units, and benchmark date. The published examples above do not establish a universal cost ranking across accelerators.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Quick Recap

Bestseller No. 2
MX3 M.2 AI Accelerator
MX3 M.2 AI Accelerator
Software and Documentation can be accessed at the MemryX developer website
$169.00
Bestseller No. 3
waveshare Hailo-8 M.2 AI Accelerator Module, Compatible with Raspberry Pi 5, Supports Linux/Windows Systems, Based On The 26TOPS Hailo-8 AI Processor, Module Only
waveshare Hailo-8 M.2 AI Accelerator Module, Compatible with Raspberry Pi 5, Supports Linux/Windows Systems, Based On The 26TOPS Hailo-8 AI Processor, Module Only
✅Scalable, enabling simultaneous processing of multi-streams & multi-models; ✅Enabling real-time, low latency and high-efficiency AI inferencing on the edge devices
$219.99
Bestseller No. 4
Tesla L40S 48GB AI HPC Graphics Accelerator
Tesla L40S 48GB AI HPC Graphics Accelerator
48GB AI graphics accelerator
$6,199.00

What commonly makes an estimate misleading?

  • Using peak throughput: It can exceed the rate that meets the service’s latency target.
  • Changing the workload between systems: Different model versions, token mixes, context lengths, precision, or serving stacks make the results difficult to compare.
  • Mixing billing units: A chip-hour price and VM-hour usage must be reconciled to the same capacity basis.
  • Ignoring idle time: A high-load benchmark can understate cost per token at low or variable traffic.
  • Comparing different cost boundaries: Accelerator-only cost is not comparable to a TCO figure that also includes hosts, power, and operations.
  • Treating a vendor example as a general price: Published benchmark costs depend on their selected workload, configuration, and measurement conditions.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.