What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
AI inference cost depends on more than how many tokens a request contains. For a hosted API, the bill may reflect separate input and output rates, model choice, caching, and service tier. For self-hosting, cost also depends on how much GPU capacity must stay available, how fully it is used, and what infrastructure and operations support it. To compare options fairly, use the same model quality, request mix, output length, and latency target.
What drives the cost of serving each request?
The main drivers are the model and hardware, input and output length, context size, batching and concurrency, utilization, latency and availability requirements, and power. They interact: the same token count can require different resources depending on whether those tokens are prompt input or generated output, how much context is active, and how requests are grouped.
Model and serving hardware
A larger or more demanding model may need more accelerator memory or compute. But a GPU’s hourly price alone does not reveal the cost of producing a useful response. What matters is the throughput it achieves for the target model and request shape while meeting the required quality and latency. NVIDIA frames inference cost as an end-to-end measure involving GPUs, CPUs, networking, software, and ecosystem. That is useful vendor positioning, not independent proof that a particular accelerator is cheapest. NVIDIA’s inference materials put it this way: “Only looking at compute pricing or FLOPs per dollar gives an incomplete view of inference TCO.”
Input, output, and context length
Hosted services commonly meter prompt input and generated output separately. A long prompt takes work to process and creates a larger context that the serving system must retain during generation; a long answer requires more generation work. Therefore, two requests with the same total number of tokens do not necessarily consume the same resources or incur the same charge.
#1 Best Overall
- NVIDIA Volta GV100 Architecture — 4,608 CUDA Cores, 640 1st-Gen Tensor Cores delivering 14 TFLOPS FP32 and 112 TFLOPS deep learning performance for AI training, inference, HPC, and scientific computing workloads
- 32GB HBM2 ECC Memory — 900 GB/s Bandwidth — High-bandwidth memory on a 4096-bit bus with ECC error correction provides the memory capacity and throughput required for the largest AI models, simulations, and datasets
- PCIe 3.0 x16 Interface — 250W TDP — Standard PCIe Gen3 connectivity with passive cooling designed for enterprise rack server deployment in HPE ProLiant, Dell PowerEdge, and Supermicro platforms with adequate chassis airflow
- NVLink — Scale to 96GB Unified Memory — Connect two V100 GPUs via NVLink at 300 GB/s bi-directional bandwidth to scale GPU memory from 32GB to 96GB for larger AI training and HPC workloads
- Multi-Precision Computing — Supports FP64 (7 TFLOPS), FP32 (14 TFLOPS), FP16 (112 TFLOPS) and INT8 precision modes for flexible deployment across training, inference, and scientific simulation workloads
The KV cache holds attention state used to generate later tokens. In the specific models and systems studied in Microsoft Research’s Splitwise paper, each active generated token accesses the cache for the context so far. The paper describes prompt batching as compute-bound and token generation as limited by memory capacity in its setup. Long contexts and concurrent sequences can consequently make memory capacity and bandwidth important to throughput, though the paper’s findings are not guarantees for every model or serving system.
Batching, utilization, and latency
Batching lets a serving system handle more inference work on a hardware allocation, potentially spreading fixed serving work across more tokens. But larger batches can affect latency and memory use, and long contexts or a varied workload can constrain how much batching is practical. A service optimized for quick responses may need capacity ready before traffic arrives; work that tolerates scheduling delays may have more flexibility.
Utilization matters especially when capacity is reserved. A GPU can be inexpensive per active computation yet costly per token if it sits idle between bursts. CNCF’s OpenCost discussion distinguishes usage-based cost—attributed infrastructure for active inference—from allocation-based cost, which includes capacity and shared services needed to keep the model available. In its explanatory example, $1.00 usage-based cost and $4.00 allocation-based cost per million tokens imply 25% utilization. That is an illustration, not an industry benchmark; CNCF identifies allocation-based cost as the more relevant measure for a build-versus-buy decision. Its article notes: “Both metrics can also be expressed as cost per million tokens, but they answer different questions and should never be confused.” Read CNCF’s OpenCost explanation.
Power and energy
Power capacity and electricity affect data-center economics, but there is no single supported electricity cost for an AI request. A useful estimate would need to specify the hardware, its power draw, utilization, facility overhead, and electricity price. Microsoft Research discusses peak power draw as a direct data-center cost factor in its systems context.
Do these 3 things before closing this tab:
1Scan for outdated or missing drivers - takes under a minute2Repair Windows errors before they cause bigger problems3Fix the driver behind crashes, sound loss and screen glitchesRank #2
- Powered by Radeon AI PRO R9700 - Supercharge you workflow with the cutting-edge RDNA 4 Architecture and 2nd-gen AI Accelerators.
- 32GB GDDR6 with 256-bit memory bus - Tackle larger, more complex projects without limits.
- PCIe Gen 5 - Unlock lightning-fast data transfers with PCIe Gen 5 support.
- GIGABYTE TURBO Fan Cooling System - Indented metal cover and blower fan increase airflow intake, while the vapor chamber, all copper heat sink, and metal frame offer efficient heat dissipation. Optimized airflow design allows for easy multi-GPU scalability.
- Double Ball Bearing Fan - Delivers superior heat resistance and rotational efficiency for better performance and a longer lifespan compared to conventional sleeve fans.
A 2026 arXiv study illustrates why energy figures need their workload attached. For Llama-3.2-1B on H200 at batch 16 and context length 4K, the authors report that increasing output length from 10 to 512 tokens reduced energy per token from 7.46 to 0.72 joules, while total energy for the batched inference window rose from 1.19 to 5.93 kilojoules. In that experiment, more output tokens spread energy across a larger token count even as total energy increased; the result is not a general prediction of request cost. See the study.
How much does AI inference cost per request?
There is no universal per-request price. For a hosted API, a useful starting estimate is:
Request charge ≈ input tokens × input rate + output tokens × output rate + applicable cache, tool, or service-tier charges.
Check the provider’s current rate table for the exact model and metering rules. Some providers distinguish cache reads and writes, long-context tiers, batch or fast modes, and image or audio units. DigitalOcean’s inference pricing page is one provider-specific example that lists model-specific per-million-token rates alongside dedicated GPU-hour prices. It also states that batch inference can receive discounts of up to 50% for OpenAI and Anthropic models. Rates and offers can change; the stated maximum is not a universal discount or a guarantee for every request.
Recommended Free Tools
Rank #3
- [Local AI Inference & 70B Model Ready] Equipped with the AMD Ryzen 7 PRO 8845HS processor, NEXUS is engineered for heavy local AI workloads. With a full-size GPU bay, it runs 70B LLMs natively without an internet connection. Ideal for AI developers and tech enthusiasts who need private environment for coding and model testing.
- [132TB Mass Storage with ZFS Integrity] Features a hybrid storage architecture (3×NVMe + 4×3.5" HDD) supporting up to 132TB. Utilizing the enterprise-grade ZFS file system and ECC memory, it prevents data corruption and bit rot—a must-have for professional photographers and video editors safeguarding 4K/8K RAW footage.
- [OpenClaw-Driven Automation Workflow] The built-in OpenClaw execution layer allows complex automated tasks to be processed locally. Even when offline, your backup schedules and AI file organization continue seamlessly. Say goodbye to monthly cloud subscriptions and high latency.
- [Dual 10GbE & USB4 Ultra-Connectivity] Experience server-class speeds with dual 10GbE ports and a 40Gbps USB4 interface. It enables multi-user real-time collaboration on large project files directly from the NAS, ensuring zero-lag editing for creative studios and production teams.
- [Open-Source ZimaOS for Total Privacy] Running on the fully open-source ZimaOS, NEXUS ensures your data stays physically on-premise with no backdoors. It acts as a "Digital Fortress" for privacy-conscious families and small businesses who demand absolute data sovereignty.
To estimate a workload rather than one request, multiply representative request counts by their estimated charges, using realistic input and output lengths and any applicable rate distinctions. Include the distribution of request sizes, not just an average token count: a small number of long-context or long-output requests can affect both the bill and the capacity needed to serve traffic.
Why are input and output tokens priced differently?
Input and output tokens correspond to different phases of inference. The system processes the prompt to build the context, then generates output token by token while retaining and consulting attention state. Those phases use compute and memory differently, which is one reason providers may set different rates for input and output. A longer output also keeps generation resources active for longer; a long prompt can increase prompt-processing work and the memory footprint of subsequent generation.
The exact price difference is provider- and model-specific, not a universal ratio. Compare the current input and output rates for the same model, and account for cache or context-tier pricing where offered.
How do the main inference options differ?
| Approach | Billing or cost basis | Questions to ask |
|---|---|---|
| Managed inference API | Often input and output usage rates; may also distinguish cache, service tier, batch, or other units. | Which model and rate tier apply? What are typical input and output lengths? Are cache, tools, or service features separately charged? |
| Dedicated cloud inference | GPU-hour or instance time, potentially plus platform costs. | How much capacity must stay available? What utilization is realistic? What latency and scaling behavior are required? |
| Self-hosted infrastructure | Amortized or rented hardware plus operations and idle or allocated capacity. | What is the fully allocated cost per token at observed traffic? Which staffing, networking, storage, and resilience costs belong in the calculation? |
DigitalOcean’s pricing page documents both serverless model rates and dedicated GPU-hour prices, illustrating that the meters differ; the GPU-hour rate cannot be compared directly with an API token rate without measuring throughput and utilization for the intended model and workload.
Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Scan for outdated or missing drivers - takes under a minuteDriver Scan →Rank #4
- Professional AI & Creator Workstation: AMD Radeon AI PRO R9700 GPU with 32GB GDDR6 is engineered for AI development, professional content creation, and compute-intensive workloads.
- Massive 32GB Memory Capacity: 32GB of GDDR6 memory on a 256-bit bus provides ample bandwidth for large AI models, 8K video editing, and complex 3D rendering.
- Advanced RDNA 4 with AI Accelerators: 64 Compute Units with 3rd Gen Ray Tracing and dedicated 2nd Gen AI Accelerators for groundbreaking AI performance and visual computing.
- Professional Blower Cooling: Efficient single blower design exhausts heat directly out of the chassis, ideal for multi-GPU workstation and server configurations.
- Enterprise-Grade Thermal Solution: Vapor chamber heatsink with industrial Honeywell PTM7950 thermal interface material ensures reliable cooling under sustained professional loads.
Is self-hosted inference cheaper than an API?
It can be, but a comparison based only on GPU purchase or rental price is incomplete. Calculate the cost of keeping the model available and divide it by the work actually served. Include reserved capacity and idle time, as well as shared infrastructure and operations that support inference. Usage-based cost can describe active work; allocation-based cost better captures the bill for maintaining service at the observed traffic level.
Compare that fully allocated self-hosting cost with the API price for the same model quality, request pattern, output length, and latency requirement. A low-volume or bursty workload can leave reserved hardware underused, while a steady workload may spread capacity costs over more requests. Batching may improve efficiency, but only to the extent the workload can tolerate its latency and memory constraints.
Performance-per-dollar claims also need their scope and date. Google Cloud’s 2023 blog reported 1.7x–3.9x relative performance improvement for specified H100/A3 workloads over A2, and up to 1.8x performance-per-dollar for a specified L4 comparison. Google explicitly said its derived performance-per-dollar measure was not an official MLPerf metric and was not verified by MLCommons; the figures and prices are historical benchmark examples, not current purchasing guidance. Google Cloud’s benchmark discussion explains the scope.
Quick Recap
How to estimate your own cost
- Define the workload. Record the model, expected input and output token distributions, context lengths, concurrency, traffic peaks, and required response latency.
- Estimate hosted charges. Apply the provider’s current input and output rates and add any relevant cache, tool, modality, or service-tier charges. Recheck rates before making a decision.
- Estimate required capacity. For dedicated or self-hosted inference, determine how much hardware must be available to meet peak demand and latency targets, rather than assuming every GPU-hour is spent processing requests.
- Allocate the full cost. Include hardware or rental, idle and reserved capacity, shared infrastructure, and operations. Divide by the tokens or requests actually served over the same period.
- Compare like with like. Evaluate both approaches against the same model quality, request mix, output length, and latency target. Test batching assumptions against the actual workload rather than applying a generic efficiency or discount figure.
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

