What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

There is no single best machine for local LLMs in 2026. The right system depends on three things: whether the model and its working memory fit in the memory a machine can actually give the model, how quickly that memory can stream weights during generation, and whether your runtime runs well on that hardware.

In broad terms, NVIDIA discrete GPUs offer strong throughput when the model and its working memory fit in VRAM and your software supports the format. Apple Silicon Macs with large unified memory can hold bigger models in a compact box, and the M4 Max and M3 Ultra configurations discussed here report the highest bandwidth in this comparison. AMD’s Ryzen AI Max+ 395 (Strix Halo) offers 128 GB of unified memory at lower reported bandwidth than the two Apple configurations. Below, each figure is tied to its source, and the one-formula comparison is shown with its assumptions.

Fit comes first, and fit is not speed

Capacity is the first gate. A fit calculation says whether the weights may load under stated assumptions. It does not promise that the result will be comfortable to use. LLMHardware.io makes this distinction explicitly: its largest-model column is a capacity ceiling, and its dense Q4_K_M estimate includes an overhead allowance rather than a speed estimate (LLMHardware.io, GPU and Apple Silicon comparison for local LLMs).

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Four things that consume memory

  • Weights. A rough weight size is parameter count × bits per weight ÷ 8, before format overhead. Quantization lowers bits per weight, so a 4-bit build of a model needs far less storage than its 16-bit original.
  • Context and KV cache. The key-value cache grows with the context length you use and sits on top of the weights.
  • Runtime allocations. The inference engine reserves its own working memory.
  • Model format and backend. The file format and the backend the runtime uses both change how much memory a model needs in practice.

A fit check to run before buying

  1. Take the size of the exact quantized file you plan to run.
  2. Add the KV cache for the context length you actually use.
  3. Add runtime overhead. The Q4_K_M estimate in the comparison above includes an allowance; your runtime may need more.
  4. Compare the total against the memory the model can actually use. On a discrete GPU, that is VRAM. On a unified-memory machine, it is the shared pool, and the amount the GPU can use depends on the operating system and runtime, so check what your setup reports rather than the headline memory figure.

If the total does not fit, you have three options: a smaller model, a more compressed quantization, or offloading part of the model to system memory. Offloading lets a model run that would otherwise fail to load, but it changes the speed picture, so test it separately.

#1 Best Overall
Sale
ASUS Dual GeForce RTX 5060 Ti 16GB GDDR7 OC Edition Gaming Graphics Card
  • AI Performance: 767 AI TOPS
  • OC mode: 2632 MHz (OC mode)/ 2602 MHz (Default mode)
  • Powered by the NVIDIA Blackwell architecture and DLSS 4
  • Axial-tech fan design features a smaller fan hub that facilitates longer blades and a barrier ring that increases downward air pressure
  • A 2.5-slot design maximizes compatibility and cooling efficiency for superior performance in small chassis

Why generation speed follows memory bandwidth

Once a model is loaded, generating each token means reading the model’s weights from memory. Tom’s Hardware put it this way in its July 30, 2026 review of the Mac Studio and M4 Max (Tom’s Hardware, Exploring Apple Silicon’s local AI performance with the Mac Studio and M4 Max):

“Because the amount of computation required for each individual token at each layer is tiny, the speed of the entire decode process basically becomes dependent on how fast those model weights can be streamed in from GPU memory.”

— Jeffrey Kampman, Senior Analyst, Graphics, Tom’s Hardware, July 30, 2026

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Two consequences follow. For a model that is already loaded, memory bandwidth rather than raw compute sets the ceiling on generation speed, so a machine with more bandwidth should generate faster, all else equal. The same review’s headline reports the M4 Max ahead of the GB10 and Strix Halo systems in decode throughput while also arguing that memory bandwidth isn’t everything. Software, configuration and the rest of the pipeline account for the difference.

Rank #2
GIGABYTE GeForce RTX 5070 Ti Gaming OC 16G Graphics Card, 16GB 256-bit GDDR7, PCIe 5.0, WINDFORCE Cooling System, GV-N507TGAMING OC-16GD Video Card
  • Powered by the NVIDIA Blackwell architecture and DLSS 4
  • Powered by GeForce RTX 5070 Ti
  • Integrated with 16GB GDDR7 256bit memory interface
  • PCIe 5.0
  • WINDFORCE cooling system

The three platforms and what each one trades

The table uses the configurations reported in Tom’s Hardware’s July 30, 2026 review. The figures describe those configurations; they are not a guarantee that each one stays on sale.

System Memory as reported Reported memory bandwidth Notes
NVIDIA GB10 128 GB unified LPDDR5X 273 GB/s A unified-memory NVIDIA system, not a discrete card; compared in the July 30, 2026 review.
AMD Ryzen AI Max+ 395 (Strix Halo) 128 GB unified 256 GB/s Memory and bandwidth depend on the OEM system.
Apple M4 Max, 128 GB configuration 128 GB unified 546 GB/s The reviewed unit had 128 GB supplied for testing. At review time, the M4 Max configuration available for purchase topped out at 64 GB, with long lead times.
Apple Mac Studio M3 Ultra Not stated in the review 819 GB/s Bandwidth as reported in the July 30, 2026 review.
NVIDIA RTX 5090 (discrete) VRAM not stated in the cited sources Not stated in the cited sources Tested in the Javat and Kazakov study covered below. Check NVIDIA’s current specification sheet for VRAM.

NVIDIA discrete GPUs

A discrete GPU has a fixed VRAM pool. When the model and working memory fit, its parallel throughput is the main strength. When they do not, you are back to offloading or a smaller model, and there is no larger shared pool to spill into. The GB10 row above is a different NVIDIA design, a unified-memory system with its own bandwidth figure. LLMHardware.io also lists the RTX 4090 and Radeon RX 7900 XTX as candidates, but the material cited here gives no VRAM or bandwidth values for them, so they do not appear in the formula table.

Apple Silicon unified memory

On Apple Silicon, the GPU can access a large shared memory pool, which is why these systems are the route to bigger models in a compact machine. The reported M4 Max and M3 Ultra bandwidths (546 GB/s and 819 GB/s) are the highest in the table. Check the chip, memory size and bandwidth of the configuration you can actually buy. Apple’s Metal backend is the software path for these machines in the sources cited here.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

AMD Ryzen AI Max+ 395 (Strix Halo)

Strix Halo is a high-capacity integrated design: 128 GB of unified memory and 256 GB/s of reported bandwidth in this comparison. For this class of machine, fitting a large model often matters more than maximum decode speed. Confirm the memory and bandwidth of the specific machine, and confirm that your runtime supports the AMD path you need, such as ROCm or Vulkan. AMD’s Ryzen AI product page describes the processor family, not the memory or bandwidth of any particular OEM system, and the cited comparison does not justify treating every AMD implementation as equivalent.

Rank #3
Sale
GIGABYTE GeForce RTX 5060 WINDFORCE OC 8G Graphics Card, Cooling System, 8GB 128-bit GDDR7, PCIe 5.0, Manufactured by NVIDIA, DisplayPort & HDMI - Video Output Interface, GV-N5060WF2OC-8GD Video Card
  • Powered by the NVIDIA Blackwell architecture and DLSS 4
  • Powered by GeForce RTX 5060
  • Integrated with 8GB GDDR7 128bit memory interface
  • PCIe 5.0
  • WINDFORCE cooling system

One formula for every row: what it computes

The formula as reported

The formula used to put every row on one scale comes from a Macyou comparison page under the same title as this article (Macyou, Best Hardware for Local LLMs in 2026: Mac vs NVIDIA vs AMD). As the indexed text describes it:

seconds per token = weight size ÷ (bandwidth × 0.9075) + 3.3 ms

In words, the model’s weight size is divided by bandwidth scaled by 0.9075, and 3.3 milliseconds of fixed overhead is added per token. The reciprocal gives tokens per second. The same page acknowledges that applying the same per-token cost to CUDA and ROCm is an assumption it has not verified by measurement. We could not open that page to check how the 0.9075 coefficient or the 3.3 ms overhead was derived, so treat both as that page’s stated assumptions rather than audited methodology.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Worked estimates

To keep the rows comparable, the formula stays constant and only the bandwidth input varies, along with two round weight sizes: 20 GB and 40 GB. These weight sizes are illustrative, not a named model at a named quantization. The outputs are modeled decode rates, not observed tokens per second.

Rank #4
ASUS TUF Gaming GeForce RTX™ 5080 16GB GDDR7 OC Edition Graphics Card
  • Powered by the NVIDIA Blackwell architecture and DLSS 4. System Requirements: Minimum 850W PSU with 16-pin 12V-2x6 (12VHPWR) connector required. Verify before purchasing.
  • Military-grade components deliver rock-solid power and longer lifespan for ultimate durability. Compatibility: 348mm (13.7") length, 3.6 slots, 4.3 lbs. Confirm case clearance and slot spacing. GPU bracket included.
  • Protective PCB coating helps protect against short circuits caused by moisture, dust, or debris
  • 3.6-slot design with massive fin array optimized for airflow from three Axial-tech fans
  • Phase-change GPU thermal pad helps ensure optimal thermal performance and longevity, outlasting traditional thermal paste for graphics cards under heavy loads
System Bandwidth input 20 GB weights: modeled ms per token 20 GB weights: modeled tokens/s 40 GB weights: modeled ms per token 40 GB weights: modeled tokens/s
NVIDIA GB10 273 GB/s 84.0 11.9 164.8 6.1
AMD Ryzen AI Max+ 395 256 GB/s 89.4 11.2 175.5 5.7
Apple M4 Max (128 GB configuration) 546 GB/s 43.7 22.9 84.0 11.9
Apple Mac Studio M3 Ultra 819 GB/s 30.2 33.1 57.1 17.5
NVIDIA RTX 5090 (discrete) Not stated in the cited sources Not stated Not stated Not stated Not stated

Within this formula, the bandwidth ratio sets the ordering: the M3 Ultra input (819 GB/s) gives roughly 2.8 to 2.9 times the modeled rate of the GB10 input (273 GB/s) at either weight size. The discrete RTX 5090 row cannot be filled, because the cited sources do not give its bandwidth, and applying this formula to a discrete card is itself an assumption.

Methodology for every row:

  • Model and quantization: none. The 20 GB and 40 GB values are round weight sizes, not a named model.
  • Weight-size convention: total weight size in GB, used directly as the numerator.
  • Bandwidth source: Tom’s Hardware review, July 30, 2026, for the four Apple, NVIDIA and AMD rows.
  • Fixed overhead: 3.3 ms per token, as stated in the formula.
  • Memory-fit condition: both weight sizes are assumed to fit in the memory the system can give the model. That is not verified for any particular machine, and the rows ignore KV cache and runtime overhead.
  • Excluded factors: prompt processing, batch size and concurrency, context-length growth, power and thermal limits, and backend-specific kernels. Because the formula is bandwidth-based, it models decode only.

What the formula can and cannot tell you

  • It can order rows that share the same weight size, bandwidth source and overhead, which is the only purpose it serves in this article.
  • It cannot predict prompt processing, concurrent serving, or the speed of a discrete GPU whose bandwidth is not in the cited sources.
  • It cannot confirm that a model fits. Run the fit check first, then use the formula only for relative decode rates.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What measured results add

The Javat and Kazakov study, Silicon Showdown: Performance, Efficiency, and Ecosystem Barriers in Consumer-Grade LLM Inference (arXiv, May 1, 2026), supplies the measured figures in this article. Each applies only to its own test.

  • 1.6× throughput from NVFP4: on an RTX 5090 running TensorRT-LLM in the authors’ test, NVFP4 reached 151 tokens per second against 92 tokens per second for optimized BF16. That is one experiment on one GPU and one runtime, not a general RTX 5090 figure. It also shows that the number format a model runs in can change throughput on the same card.
  • 23× energy-efficiency advantage: the authors report that the Apple M3 Ultra was 23 times more energy-efficient than the RTX 5090 in their lightweight 1.5B-model baseline. The result does not extend to larger models or other workloads.

The authors also describe backend and memory constraints on the platforms they test, a reminder that the software stack counts alongside the silicon. No cross-vendor measurement under one protocol is reported in the sources for this article, so the modeled table is the only place all rows sit on one scale, and it is a model.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Quick Recap

SaleBestseller No. 1
ASUS Dual GeForce RTX 5060 Ti 16GB GDDR7 OC Edition Gaming Graphics Card
ASUS Dual GeForce RTX 5060 Ti 16GB GDDR7 OC Edition Gaming Graphics Card
AI Performance: 767 AI TOPS; OC mode: 2632 MHz (OC mode)/ 2602 MHz (Default mode); Powered by the NVIDIA Blackwell architecture and DLSS 4
$786.37
Bestseller No. 2
GIGABYTE GeForce RTX 5070 Ti Gaming OC 16G Graphics Card, 16GB 256-bit GDDR7, PCIe 5.0, WINDFORCE Cooling System, GV-N507TGAMING OC-16GD Video Card
GIGABYTE GeForce RTX 5070 Ti Gaming OC 16G Graphics Card, 16GB 256-bit GDDR7, PCIe 5.0, WINDFORCE Cooling System, GV-N507TGAMING OC-16GD Video Card
Powered by the NVIDIA Blackwell architecture and DLSS 4; Powered by GeForce RTX 5070 Ti; Integrated with 16GB GDDR7 256bit memory interface
$1,162.49
SaleBestseller No. 3
GIGABYTE GeForce RTX 5060 WINDFORCE OC 8G Graphics Card, Cooling System, 8GB 128-bit GDDR7, PCIe 5.0, Manufactured by NVIDIA, DisplayPort & HDMI - Video Output Interface, GV-N5060WF2OC-8GD Video Card
GIGABYTE GeForce RTX 5060 WINDFORCE OC 8G Graphics Card, Cooling System, 8GB 128-bit GDDR7, PCIe 5.0, Manufactured by NVIDIA, DisplayPort & HDMI - Video Output Interface, GV-N5060WF2OC-8GD Video Card
Powered by the NVIDIA Blackwell architecture and DLSS 4; Powered by GeForce RTX 5060; Integrated with 8GB GDDR7 128bit memory interface
$459.99
Bestseller No. 4
ASUS TUF Gaming GeForce RTX™ 5080 16GB GDDR7 OC Edition Graphics Card
ASUS TUF Gaming GeForce RTX™ 5080 16GB GDDR7 OC Edition Graphics Card
3.6-slot design with massive fin array optimized for airflow from three Axial-tech fans; Auto-Extreme precision automated manufacturing helps ensure higher reliability
$1,831.31

Choosing by workload

  • One user, a model that fits in VRAM, and a supported format: an NVIDIA discrete GPU is the natural first candidate, because the throughput shown in the measured results applies in that setting.
  • The largest model you can load in one pool, in a compact machine: Apple Silicon or Strix Halo. Choose by bandwidth (546 or 819 GB/s against 256 GB/s in the tables) and by the runtime you need. Strix Halo is the more natural fit where model capacity matters more than maximum decode speed.
  • Power-sensitive use: the efficiency result above favors the Apple M3 Ultra in one small test. Treat it as a direction, not a verdict.
  • Several users at once: none of the cited sources measure concurrent serving, and the formula does not model it. Test your own serving stack.
  • Long prompts or document work: prompt processing is a separate axis from decode speed. Check measured prompt-processing results for your model and runtime before relying on a decode-only estimate.

Before you buy

  • Confirm the exact SKU, memory size and bandwidth of the configuration you can order. Reviewed units are not always the configuration on sale.
  • Recheck prices and availability. Tom’s Hardware reported differing retailer prices and configurations, and LLMHardware.io notes that its listed prices are indicative and set by retailers.
  • Confirm that your runtime supports the model format and quantization you plan to run on that backend (Metal, CUDA, ROCm or Vulkan).
  • Run the fit check with your intended context length, not just the headline weight size.
  • Benchmark the exact model, quantization, runtime and context length you plan to use. No published measurement covers every combination discussed here.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.