There is no universal price for training an AI model on a supercomputer. Estimate the workload’s runtime on the intended system, apply that system’s actual billing or allocation terms, then add storage, networking, data movement, and operational costs that are not included. A cloud quote, a paid facility commitment, and an awarded research allocation are different forms of access—not directly comparable hourly prices.
What determines the cost?
Model size alone is not enough to calculate a reliable training bill. Cost depends on the specific workload, how efficiently it runs on the chosen hardware, and how access is priced or allocated. Pretraining and fine-tuning can require very different amounts of computation, so decide which one you are estimating before comparing systems.
| # | Preview | Product | Price | |
|---|---|---|---|---|
| 1 |
|
MINISFORUM MS-02 Ultra Workstation Mini PC, Intel Core Ultra 9 285HX (24C/24T, up to 5.5GHz), PCIe... | $1,659.00 | Buy on Amazon |
| 2 |
|
GMKtec EVO-X2 AI Mini PC Ryzen Al Max+ 395 Superchip 128GB LPDDR5X 2TB SSD | $3,649.99 | Buy on Amazon |
Start by recording the model architecture and size, training objective, token or sample volume, sequence length, precision, parallelism plan, number of training steps, and number of runs. These define the work to be done; a parameter count or a GPU’s advertised peak performance does not tell you how long that work will take.
Estimate the workload in six steps
- Define the full workload. Include training, evaluation, planned experiments, and the number of runs—not just the final successful run.
- Benchmark a representative slice on the target system. Measure end-to-end throughput and wall-clock time, including data loading and communication. A small test is useful only if its configuration reflects the intended scale.
- Extrapolate cautiously. Use measured throughput and the remaining work to estimate runtime. Check that the benchmark and planned run use comparable hardware, precision, parallelism, and data handling. There is no universal conversion from parameter count or peak GPU performance to training time.
- Convert runtime to the system’s resource unit. For a cloud setup, calculate elapsed instance time for the configured resources and use the applicable machine, GPU, region, and discount terms. For a facility quoting node-hours, multiply elapsed hours by the number of nodes, using that facility’s definition of a node. Confirm whether the quote is for GPU-hours, node-hours, reserved capacity, or another unit.
- Add costs outside the compute line. Check how the system bills for training data, checkpoints and outputs, networking, data transfer, setup, and operational support.
- Document uncertainty and assumptions. Account explicitly for evaluation, checkpointing, restarts, debugging, unsuccessful runs, and any reserved-capacity or minimum-commitment requirements. These are project-specific assumptions, not a universal runtime multiplier.
For example, if a representative benchmark measures a run’s elapsed time on the planned configuration, use that elapsed time and the provider’s or facility’s stated billing unit to price the compute. Do not multiply by a supposed standard efficiency factor: the sources here establish no universal runtime conversion or failure allowance.
Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Scan for outdated or missing drivers - takes under a minute3Clear out junk files and repair common Windows errors#1 Best Overall
- High-Performance AI Processor:The MS-02 Ultra features an Intel Core Ultra 9 285HX (24C/24T, up to 5.5 GHz, 13 TOPS NPU), delivering fast and efficient performance for AI inference, algorithm development, and media workloads. A PCIe x16 expansion slot supports desktop-class GPU upgrades for advanced model training and accelerated computing tasks. It's ideal for creators, engineers, and teams handling intensive parallel workloads.
- 4 × M.2 PCIe 4.0 + 4 × DDR5 SODIMM slots:Four DDR5 SODIMM slots support up to 256 GB of memory, while ECC helps maintain data integrity in mission-critical environments. Four PCIe 4.0 M.2 slots support up to 24 TB of storage, supporting RAID 0/1/5/10, combining high-speed performance with data protection. It allows for the creation of independent scratch disks, media libraries, and project drives, providing high-throughput for production workflows.
- PCIe & USB 4.0 v2: Up to three PCIe slots can be equipped, including a dual-slot x16 GPU. The main slot supports PCIe 5.0, meeting the needs of high-bandwidth creative and computing workloads. USB 4.0 v2 (80Gbps) supports high-bandwidth external storage and displays.
- Ultra-fast Networking: Wi-Fi 7 further enhances wireless performance with next-generation speeds and low-latency stability. Intelligent bandwidth switching optimizes throughput in different network environments, ensuring optimal performance for enterprise or local networks. Dual 25GbE ports (providing up to approximately 3.125 GB/s bandwidth, about 25 times faster than traditional 1GbE), enabling seamless large-scale file transfers and parallel computing. 10GbE and 2.5GbE ports, with support for Intel vPro technology, ensure enterprise-grade remote management and deployment flexibility.
- Server-grade thermal architecture: Utilizing a dedicated CPU/GPU airflow design, equipped with a 6-pipe dual-fan cooler, it maintains stable performance even under sustained loads, delivering up to 140W Turbo power while maintaining a 100W TDP, and operating with noise levels as low as 36 dB. An integrated 350W power supply ensures stable and reliable output for demanding computing tasks and fully loaded extended configurations.
Build a project estimate, not just a GPU quote
Use the worksheet below to keep measured values, quotes, and assumptions distinct. A line item applies only if it is billed separately or represents a real project cost.
| Cost area | How to estimate it | What to verify |
|---|---|---|
| Training compute | Measured runtime × configured nodes or instances × applicable rate, or the resource units consumed | Representative benchmark, billing unit, and current provider or facility quote |
| Data preparation and evaluation | Estimate as separate measured or planned jobs | Workflow schedule and benchmark |
| Storage | Capacity × duration × applicable rate, if charged separately | Dataset size, checkpoint retention, and filesystem or object-storage terms |
| Network and data movement | Apply transfer and network charges, if billed | Data location, transfer plan, and provider terms |
| Setup and operations | Estimate project-specific labor or service costs | Staffing and selected services |
| Energy, if separately billed | Measured or estimated energy × applicable billed energy rate | Power measurement or model and the actual billing arrangement |
| Contingency | Add an explicitly stated scenario allowance | Project risk assumptions; there is no sourced universal percentage |
Total project estimate = compute + storage + networking and data movement + setup and operations + any separately billed energy + explicitly stated contingency. Not every system bills each item separately. Label each input as quoted, benchmarked, modeled, or assumed so a reader can see what is known and what remains uncertain.
Cloud pricing: check the configured system and the exclusions
Google Cloud’s GPU pricing page lists prices by region and states that GPU charges are additional to the VM machine-type charge. Its displayed GPU table excludes disk, networking, and VM instance pricing; the page directs customers to its pricing calculator to estimate configured resources. Rates, zones, and discounts can change, so use a current quote for the selected configuration rather than treating a listed GPU rate as the whole training bill.
When reviewing any cloud quote, verify whether it includes CPU and memory, attached or local storage, networking, licenses, support, taxes, and committed-capacity terms. These vary by service and contract. Google’s cost framework also identifies compute, networking, storage, training-data and adapter-layer storage, application setup, and operational support as cost areas.
The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Facility allocations and paid supercomputing access
A publicly funded allocation can reduce or eliminate a direct compute charge for an eligible project, but it is not equivalent to a retail hourly price. Eligibility, award limits, resource units, storage terms, scheduling, and access windows determine what the allocation is worth for a particular workload.
NERSC Perlmutter: an allocation, not a retail rate
In its March 18, 2026 AI for Science call, NERSC said accepted projects could initially receive up to 10,000 Perlmutter GPU node-hours, with associated filesystem storage quotas. Each Perlmutter GPU node has four A100 GPUs. Awards apply to the 2026 allocation year, which runs through January 19, 2027. This is a time-bounded research allocation example, not a cash price per hour; eligibility and the award a project receives matter.
Rank #2
- EVOLUTION RYZEN AI MAX+ 395 MINI PC - GMKtec EVO-X2 is the next evolution in AI mini PC Ryzen Strix Halo series. Thanks to AMD Simultaneous Multithreading (SMT) the core-count is effectively doubled, to 32 threads. Ryzen AI Max+ 395 has 64 MB of L3 cache and can boost up to 5.1 GHz, depending on the workload. The Ryzen AI Max+ 395 is currently rated as the "most powerful x86 APU" on the market for AI computing.
- AI NPU with XDNA 2 ARCHITECTURE - Powered by 16 “Zen 5” CPU cores, 50+ peak AI TOPS XDNA 2 NPU and a truly massive integrated GPU driven by 40 AMD RDNA 3.5 CUs, the Ryzen AI MAX+ 395 is a transformative upgrade and delivers a significant performance boost over the competition. The Ryzen AI Max+ 395 excels in consumer AI workloads like the llama.cpp-powered application: LM Studio. Shaping up to be the must-have app for client LLM workloads, LM Studio allows users to locally run the latest language model without any technical knowledge required and unleash their creativity and productivity.
- AMD RADEON 8090S iGPU GAMING PC - The AMD Radeon RX 8060S offers all 40 CUs with up to 2.9 GHz graphics clock and uses the new RDNA 3.5 architecture. The powerful iGPU is positioned between an RTX 4060 and 4070 laptop GPU and therefore enables gaming in FHD at maximum details in most demanding games. The 8060S can also utilize the full 128GB pool, which is perfect for running LLMs such as Deepseek 70B Q8, which runs comfortably on this machine.
- EIGHT CHANNEL LPDDR5X - LPDDR5X is a new ground breaking memory small form factor installed on-board. With blazing speeds up to to 8000MT/s, it runs 1.5x faster than the DDR5 SODIMMs; 90% better performance over DDR5 SODIMMs in video conferencing and photo editing; 30% better performance in productivity apps; 12% better performance in digital content workloads.
- QUAD SCREEN 8K DISPLAY SUPPORT - EVO-X2 AI Mini PC support 4-screen 4K/8K output via HDMI 2.1 (8K@60Hz), DisplayPort 1.4 (4K@60Hz), and dual USB 4 40Gbps Transfer speed (supporting PD3.0/DP1.4/DATA). Ideal for gaming, video editing, and multitasking, it provides expansive and crisp multi-display support.
OLCF Lux: paid proprietary access with a minimum commitment
Oak Ridge Leadership Computing Facility (OLCF) says its Lux system reserves half of its annual 3.5 million node-hours for the Genesis Mission; the remaining half is available for proprietary paid use under the DOE User Facility rate. OLCF states a minimum commitment of 175,000 node-hours per six months and says allocated storage access is included. The page describes more than 4,000 MI355X GPUs across 500-plus nodes. These capacity and commercial terms are time-sensitive; confirm the current offer with OLCF before using them in a quote.
For comparison, look at cash price, eligibility, the resource unit, minimum commitment, included storage, capacity availability, and scheduling or access terms. An awarded allocation has project and opportunity costs, but it should not be represented as a cash charge per hour unless the facility specifies one.
Quick wins for a faster PC:
Repair Windows errors before they cause bigger problemsFix Now →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Clear out junk files and repair common Windows errorsFree Scan →Energy estimates are not training prices
An OLCF tutorial slide titled “Energy Budget against Model Size (~P^2)” reports estimates of 284 gigajoules for a 22B-parameter model, 17.65 terajoules for a 175B-parameter model, and 662 terajoules for a 1T-parameter model. The year of the Oak Ridge National Laboratory slide is not stated. Its method uses iteration time, tokens consumed per iteration, average active power, and total MI250X GPU-card count; it says GPU-level energy was measured with rocm-smi.
These figures depend on the tutorial’s method and assumptions. They are not current prices, universal energy coefficients, or directly comparable training bills. Energy is a physical quantity; converting it to a monetary charge requires the applicable electricity or facility billing rate. In a cloud bill, energy may be embedded in the service rate rather than metered and charged separately. Add a separate power line only when the arrangement actually bills it.
Compare systems using a workload benchmark
Hardware specifications help describe what a system is, but do not substitute for measuring your workload. OLCF’s Summit system page describes a Summit node with six NVIDIA V100 GPUs and reports 13 MW peak system power consumption. That is system-specific context, not a current cloud quote or an estimate of the power consumed by an individual training run.
OLCF’s Lux page describes MI355X accelerators, high-bandwidth memory, Slurm and Kubernetes scheduling, and included access to the Orion filesystem for Lux allocations. Such details can inform a configuration comparison, but advertised peak performance alone cannot predict application throughput. Benchmark a representative slice and confirm deployment and access conditions with the facility.
Quick Recap
What to put in the final estimate
- Workload assumptions: objective, data volume, sequence length, precision, parallelism, steps, and planned runs.
- Benchmark evidence: system configuration, measured throughput, elapsed time, and whether data loading and communication were included.
- Access terms: cloud instance and GPU charges, facility resource unit, allocation eligibility, or any minimum commitment.
- Other costs: storage, network and data movement, setup, operations, and energy only if separately billed.
- Uncertainty: evaluation, restarts, debugging, unsuccessful work, and a clearly identified project-specific contingency.
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

