Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

If a local LLM is slow or runs out of memory, first find out whether the problem happens while loading the model, processing a prompt, or generating tokens. Then verify whether it is actually using the GPU and identify what is occupying memory. Only after that should you change context, concurrency, model precision, or hardware.

Why is my local LLM so slow?

“Slow” can describe three different delays, and each points to a different cause. Record them separately while keeping the model, prompt, context setting, backend, and machine the same:

  • Load and first response: time how long the model takes to load and return its first token. Repeated startup delays can occur if the model is not kept in memory.
  • Prompt processing: note the delay before the model begins answering. Long prompts and large context settings can increase work and memory use.
  • Token generation: measure how quickly the answer continues after it starts. Device placement, CPU thread count, and the chosen backend can affect this stage.

Compare one change at a time; a speed result from another GPU, model, or backend does not predict yours. The official material reviewed does not establish a universal acceptable tokens-per-second threshold.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Why is my GPU not being used?

Check the runtime’s placement diagnostics before changing settings or shopping for hardware. The selected backend and its installed GPU build must support the accelerator you intend to use.

#1 Best Overall
ASUS Dual Radeon RX 9060 XT 16GB GDDR6 Gaming Graphics Card
  • Axial-tech fans now feature a smaller fan hub that facilitates longer blades and a barrier ring that increases downward air pressure
  • 2.5-slot design allows for greater build compatibility while maintaining cooling performance
  • 0dB technology lets you enjoy light gaming in relative silence
  • Dual BIOS switch lets you toggle between Quiet and Performance BIOS profiles
  • Dual ball fan bearings last up to twice as long as sleeve bearing designs

Check placement in llama.cpp

Inspect the startup output for GPU-offloaded layers and VRAM use. The llama.cpp project documentation identifies these diagnostics as evidence of GPU use.

Check placement in Ollama

Run ollama ps and inspect the Processor field. It distinguishes a model placed entirely on GPU, entirely on CPU, or split between them. If the model is on CPU or split, check whether your installed Ollama version and GPU backend support the selected accelerator.

Rank #2
GIGABYTE GeForce RTX 5070 Ti Gaming OC 16G Graphics Card, 16GB 256-bit GDDR7, PCIe 5.0, WINDFORCE Cooling System, GV-N507TGAMING OC-16GD Video Card
  • Powered by the NVIDIA Blackwell architecture and DLSS 4
  • Powered by GeForce RTX 5070 Ti
  • Integrated with 16GB GDDR7 256bit memory interface
  • PCIe 5.0
  • WINDFORCE cooling system

How do I fix CUDA out of memory?

GPU out-of-memory means the workload needs more VRAM than is available, but model weights are only one part of that workload. Memory may also be needed for the KV cache, activations, runtime and communication buffers, adapters, multimodal state, and other allocations. Long context and concurrent requests can push a deployment over its limit even when the weights appear to fit.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

NVIDIA’s NIM troubleshooting guide gives a rough weight-memory estimate: parameter count × bytes per parameter ÷ tensor-parallelism degree. For example, it estimates approximately 16 GB of weight memory for Llama 3.1 8B in BF16 at tensor parallelism 1 (8 billion parameters × 2 bytes). That is an estimate for weights, not a fit guarantee: cache and runtime overhead still need space.

Rank #3
Sale
GIGABYTE GeForce RTX 5060 WINDFORCE OC 8G Graphics Card, Cooling System, 8GB 128-bit GDDR7, PCIe 5.0, Manufactured by NVIDIA, DisplayPort & HDMI - Video Output Interface, GV-N5060WF2OC-8GD Video Card
  • Powered by the NVIDIA Blackwell architecture and DLSS 4
  • Powered by GeForce RTX 5060
  • Integrated with 8GB GDDR7 128bit memory interface
  • PCIe 5.0
  • WINDFORCE cooling system

Identify when the error occurs

  • During weight loading: try a smaller model or a lower-precision format supported by your backend and hardware. A supported multi-GPU configuration may also distribute the model.
  • After weights load, during cache allocation: reduce the maximum context to what the task actually needs. KV-cache demand grows with context and can also grow with parallel requests.
  • During fragmentation, graph capture, or warmup: do not assume that reducing context alone will resolve it. Read the runtime’s specific error and follow its backend-specific diagnosis.

NIM flags and procedures are for NIM/vLLM configurations; do not copy them into Ollama or llama.cpp without checking that runtime’s documentation.

How can context length and concurrency reduce memory use?

Set context to the task rather than leaving an unnecessarily large limit, and keep only the concurrent requests the workload needs. Ollama’s FAQ describes a 4096-token default in the version reviewed, with context overrides available through the OLLAMA_CONTEXT_LENGTH environment variable, the CLI parameter, or the API’s num_ctx. Defaults can change, so check the documentation for your installed version.

Rank #4
Sale
GIGABYTE Radeon RX 9070 XT Gaming OC 16G Graphics Card, PCIe 5.0, 16GB GDDR6, GV-R9070XTGAMING OC-16GD Video Card
  • Powered by Radeon RX 9070 XT
  • WINDFORCE Cooling System
  • Hawk Fan
  • Server-grade Thermal Conductive Gel
  • RGB Lighting

Ollama also explains that parallel processing increases required RAM or VRAM with both the number of parallel requests and context length. Avoid loading or running more models and requests at once than the workload requires.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Ollama cache options

The current Ollama FAQ describes Flash Attention as a way to reduce memory use as context grows, and documents quantized KV-cache options when Flash Attention is enabled. Ollama says its q8_0 cache uses approximately half the memory of f16 with very small stated precision loss; q4_0 uses approximately a quarter with small-to-medium stated loss that may be more noticeable at higher context. These are Ollama’s claims, not guarantees across models or backends. Test answer quality on representative tasks before relying on a lower-precision cache.

Best Value
Sale
ASUS Prime Radeon RX 9070 XT 16GB GDDR6 OC Edition Gaming Graphics Card
  • Axial-tech fans now feature a smaller fan hub that facilitates longer blades and a barrier ring that increases downward air pressure
  • Phase-change GPU thermal pad helps ensure optimal heat transfer, lowering GPU temperatures for enhanced performance and reliability
  • 2.5-slot design allows for greater build compatibility while maintaining cooling performance
  • Dual-ball fan bearings last up to twice as long as standard conventional sleeve bearings designs
  • 0dB technology lets you enjoy light gaming in relative silence
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Which CPU, model, or backend settings should I change?

Adjust CPU threads incrementally

Using more CPU threads is not always faster. The llama.cpp project’s token-generation troubleshooting guidance warns that excessive thread counts can oversaturate the CPU. For unusually slow generation, it suggests trying one thread, then increasing gradually and backing down if performance worsens. Keep the prompt and machine constant during comparisons.

Choose model precision with quality in mind

Quantization can reduce weight memory, but the right trade-off depends on the model and task. Compare outputs on representative prompts as well as checking whether the model fits; lower memory use alone does not establish acceptable answer quality.

Match the backend and model residency to the workload

Ollama documents preloading or keeping a model in memory to avoid repeated startup time, as well as unloading it to free memory. Keeping a model resident can reduce startup delay but uses memory that may otherwise be available to other models or applications. NVIDIA’s NIM troubleshooting guide recommends choosing an inference backend in light of operating system, model format, GPU architecture and memory, API needs, and throughput target; the suitable configuration is environment-specific.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

When should I add GPU hardware?

Consider more VRAM only after placement checks and memory diagnosis show that the model and useful context do not fit the available device. First compare whether a smaller model, supported lower precision, shorter context, or fewer concurrent requests would meet the need. Hardware capacity alone does not determine speed or compatibility.

Before choosing a GPU or backend, compare the whole workload:

Quick Recap

Bestseller No. 1
ASUS Dual Radeon RX 9060 XT 16GB GDDR6 Gaming Graphics Card
ASUS Dual Radeon RX 9060 XT 16GB GDDR6 Gaming Graphics Card
0dB technology lets you enjoy light gaming in relative silence; Dual BIOS switch lets you toggle between Quiet and Performance BIOS profiles
$529.00
Bestseller No. 2
GIGABYTE GeForce RTX 5070 Ti Gaming OC 16G Graphics Card, 16GB 256-bit GDDR7, PCIe 5.0, WINDFORCE Cooling System, GV-N507TGAMING OC-16GD Video Card
GIGABYTE GeForce RTX 5070 Ti Gaming OC 16G Graphics Card, 16GB 256-bit GDDR7, PCIe 5.0, WINDFORCE Cooling System, GV-N507TGAMING OC-16GD Video Card
Powered by the NVIDIA Blackwell architecture and DLSS 4; Powered by GeForce RTX 5070 Ti; Integrated with 16GB GDDR7 256bit memory interface
$1,162.49
SaleBestseller No. 3
GIGABYTE GeForce RTX 5060 WINDFORCE OC 8G Graphics Card, Cooling System, 8GB 128-bit GDDR7, PCIe 5.0, Manufactured by NVIDIA, DisplayPort & HDMI - Video Output Interface, GV-N5060WF2OC-8GD Video Card
GIGABYTE GeForce RTX 5060 WINDFORCE OC 8G Graphics Card, Cooling System, 8GB 128-bit GDDR7, PCIe 5.0, Manufactured by NVIDIA, DisplayPort & HDMI - Video Output Interface, GV-N5060WF2OC-8GD Video Card
Powered by the NVIDIA Blackwell architecture and DLSS 4; Powered by GeForce RTX 5060; Integrated with 8GB GDDR7 128bit memory interface
$459.99
SaleBestseller No. 4
GIGABYTE Radeon RX 9070 XT Gaming OC 16G Graphics Card, PCIe 5.0, 16GB GDDR6, GV-R9070XTGAMING OC-16GD Video Card
GIGABYTE Radeon RX 9070 XT Gaming OC 16G Graphics Card, PCIe 5.0, 16GB GDDR6, GV-R9070XTGAMING OC-16GD Video Card
Powered by Radeon RX 9070 XT; WINDFORCE Cooling System; Hawk Fan; Server-grade Thermal Conductive Gel
$814.28
SaleBestseller No. 5
ASUS Prime Radeon RX 9070 XT 16GB GDDR6 OC Edition Gaming Graphics Card
ASUS Prime Radeon RX 9070 XT 16GB GDDR6 OC Edition Gaming Graphics Card
0dB technology lets you enjoy light gaming in relative silence; Dual BIOS switch lets you toggle between Quiet and Performance BIOS profiles
$829.00
Decision axis What to compare
Memory fit Weight size and precision, intended context and KV cache, plus runtime headroom.
Latency and throughput Load/first-token time, prompt-processing delay, and generation speed on your machine with the actual workload.
Output quality Quantized model and cache answers on representative prompts; quality impact depends on model and task.
Compatibility Operating system, GPU architecture, model format, backend, API needs, and supported precision.
Operational trade-off Concurrency and model residency versus memory needed by other models and applications.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.