Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Find out which part of local inference is slow before changing settings: model loading, prompt processing, or token generation. Then check the runtime’s logs and memory use to see whether the cause is GPU placement, context and KV-cache allocation, CPU threading, or a model that exceeds available memory. Ollama, llama.cpp, and vLLM use different controls, so verify commands against the version you have installed.

Identify which part of inference is slow

“Slow inference” can describe three different delays. Separate them before tuning, because a fix for model loading will not necessarily improve token generation.

  • Model loading: the delay before the runtime says the model is ready.
  • Prompt processing: the delay before the first generated token, while the model reads the input.
  • Token generation: the rate at which new tokens appear after generation starts.

If model loading is slow

For vLLM, model downloads depend on network conditions. Large model files, shared or network filesystems, and host-memory pressure can also slow loading; swapping to disk can make it especially sluggish. Check whether the model is already local, whether its files are on a local disk, and whether system memory is under pressure. The vLLM v0.18.2 troubleshooting guide suggests using a local model path and local disk, monitoring CPU memory, and trying --load-format dummy to isolate loading behavior.

If the first token is slow but later tokens are not

Look at prompt processing and context size. A long prompt requires more processing before generation begins, and the selected context length also affects memory allocation. If the runtime exposes separate prompt and generation metrics, compare them rather than treating both as one speed number.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
ASUS Dual Radeon RX 9060 XT 16GB GDDR6 Gaming Graphics Card
  • Axial-tech fans now feature a smaller fan hub that facilitates longer blades and a barrier ring that increases downward air pressure
  • 2.5-slot design allows for greater build compatibility while maintaining cooling performance
  • 0dB technology lets you enjoy light gaming in relative silence
  • Dual BIOS switch lets you toggle between Quiet and Performance BIOS profiles
  • Dual ball fan bearings last up to twice as long as sleeve bearing designs

If every generated token is slow

Check whether the model is actually running on the GPU, whether some layers remain on the CPU, and whether CPU threads or temporary debugging instrumentation are limiting performance. The sections below explain how to verify those conditions.

Check whether the GPU is doing inference work

Seeing a GPU in the system does not prove the runtime is using it. Confirm actual placement in startup logs. In CUDA-enabled llama.cpp runs, look for diagnostics showing how many layers were offloaded and how much VRAM they use. The llama.cpp performance guide identifies these startup messages as evidence of GPU use.

In llama.cpp, -ngl (also available as --gpu-layers) requests GPU offload. A large value asks the runtime to offload as many layers as fit; it does not guarantee that every layer will fit. If logs show partial offload, CPU-resident layers may constrain speed. Confirm that your installed build supports the device backend, then check the startup output rather than assuming a flag succeeded.

Rank #2
GIGABYTE GeForce RTX 5070 Ti Gaming OC 16G Graphics Card, 16GB 256-bit GDDR7, PCIe 5.0, WINDFORCE Cooling System, GV-N507TGAMING OC-16GD Video Card
  • Powered by the NVIDIA Blackwell architecture and DLSS 4
  • Powered by GeForce RTX 5070 Ti
  • Integrated with 16GB GDDR7 256bit memory interface
  • PCIe 5.0
  • WINDFORCE cooling system

The llama.cpp server CLI reference documents --fit, enabled by default, as a setting that adjusts unspecified arguments to fit device memory. It also describes multi-GPU placement options: layer split (the documented default), row split, and experimental tensor split. These modes distribute work differently; use the current CLI reference for your build before changing them. Flags for llama.cpp should not be copied to Ollama or vLLM.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Reduce context and KV-cache memory carefully

Context length is not just a prompt-capacity setting: the KV cache used to retain context consumes memory. If you set a much longer context than your task needs, reduce it to a realistic length first. This can relieve memory pressure, though it also limits how much conversation or source text the model can retain.

Ollama Flash Attention

Ollama documents Flash Attention as a way to significantly reduce memory use as context grows when the selected backend and devices support it. To force it, set OLLAMA_FLASH_ATTENTION=1; set it to 0 to disable it. See the Ollama FAQ for the documented conditions and settings.

Rank #3
Sale
GIGABYTE GeForce RTX 5060 WINDFORCE OC 8G Graphics Card, Cooling System, 8GB 128-bit GDDR7, PCIe 5.0, Manufactured by NVIDIA, DisplayPort & HDMI - Video Output Interface, GV-N5060WF2OC-8GD Video Card
  • Powered by the NVIDIA Blackwell architecture and DLSS 4
  • Powered by GeForce RTX 5060
  • Integrated with 8GB GDDR7 128bit memory interface
  • PCIe 5.0
  • WINDFORCE cooling system

Ollama KV-cache precision

With Flash Attention enabled, Ollama documents the OLLAMA_KV_CACHE_TYPE setting. Its FAQ describes these approximate memory and quality tradeoffs:

Ollama cache type Documented memory use relative to f16 Documented quality tradeoff
f16 Default; reference value Reference precision
q8_0 Approximately half of f16 Very small loss
q4_0 Approximately one quarter of f16 Small-to-medium loss, potentially more noticeable at higher context sizes

These are Ollama’s approximate, project-documented comparisons, not guaranteed measurements for every model or runtime. Ollama says the response-quality effect depends on the model and task; models with a high GQA count may be more sensitive to reduced precision. Try a cache setting on representative prompts and check output quality rather than assuming the smallest cache is best.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Diagnose out-of-memory errors

An OOM error means a runtime allocation could not be satisfied, but the message alone does not identify what consumed the memory. Check runtime logs and resource use to distinguish among model weights, KV cache, concurrency, and other runtime allocations. Also identify whether the shortage is in GPU memory or host RAM: the remedies differ.

Rank #4
Sale
GIGABYTE Radeon RX 9070 XT Gaming OC 16G Graphics Card, PCIe 5.0, 16GB GDDR6, GV-R9070XTGAMING OC-16GD Video Card
  • Powered by Radeon RX 9070 XT
  • WINDFORCE Cooling System
  • Hawk Fan
  • Server-grade Thermal Conductive Gel
  • RGB Lighting

vLLM states that “If the model is too large to fit in a single GPU, you will get an out-of-memory (OOM) error.” Its troubleshooting documentation points users toward memory-reduction options. There is no universal VRAM threshold for local LLMs: fit depends on the model, its representation, context and runtime allocation, and the runtime’s placement options.

Try lower-impact changes first

  1. Reduce unnecessary context and concurrency. Use only the context your task requires, and reduce simultaneous requests if the runtime’s memory use rises with concurrent work.
  2. Reduce KV-cache use where supported. For Ollama, consider the documented Flash Attention and cache-precision settings, then validate response quality.
  3. Choose a smaller or lower-memory model representation. Use a model or quantization supported by your chosen runtime. This changes the model representation and may affect output quality; compare results on the work you actually do.
  4. Adjust supported GPU placement or splitting. For llama.cpp, inspect GPU offload and consult the installed build’s options, including its fit and multi-GPU controls. Do not assume another runtime accepts the same flags.
  5. Add hardware capacity only after identifying the constrained resource. More GPU memory may help when weights or cache do not fit, but it is not a diagnosis or a universal fix; first establish the model, context, runtime, and workload involved.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Tune CPU threads without oversaturating the processor

If llama.cpp is doing CPU work, or CPU threads are part of a mixed CPU/GPU run, too many threads can oversaturate the processor and reduce performance. The llama.cpp performance guide suggests starting with one thread and doubling the count until a bottleneck appears, then scaling back. If one thread helps, it also suggests explicitly trying the number of physical CPU cores. Treat this as a troubleshooting heuristic, not a universally optimal setting.

The same guide reports a project benchmark of 9.1 tokens per second for a specific setup: an NVIDIA A6000 with 48 GB VRAM, a CPU with seven physical cores, 32 GB RAM, and a 30B-parameter Q4_0 GGML model. In its listed results, -t 4 with the stated large GPU-layer setting measured 9.1 tokens per second, while -t 7 with that setting measured 8.7 tokens per second. The page does not state a publication year. These figures describe that configuration, not expected performance on a different machine or with a current model format.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Best Value
Sale
ASUS Prime Radeon RX 9070 XT 16GB GDDR6 OC Edition Gaming Graphics Card
  • Axial-tech fans now feature a smaller fan hub that facilitates longer blades and a barrier ring that increases downward air pressure
  • Phase-change GPU thermal pad helps ensure optimal heat transfer, lowering GPU temperatures for enhanced performance and reliability
  • 2.5-slot design allows for greater build compatibility while maintaining cooling performance
  • Dual-ball fan bearings last up to twice as long as standard conventional sleeve bearings designs
  • 0dB technology lets you enjoy light gaming in relative silence

Turn off temporary debugging instrumentation

Debug settings can distort performance after the original issue has been diagnosed. vLLM warns that VLLM_TRACE_FUNCTION=1 slows token generation by over 100x and should not be used unless absolutely needed. Turn off that setting and other debugging environment variables that you enabled for troubleshooting, then measure again. See the vLLM troubleshooting guide.

Measure the change, not just the impression

Use the same model, prompt, context, and workload when comparing settings; otherwise, a faster result may simply reflect less work. The llama.cpp server reference exposes separate metrics, including llamacpp:prompt_tokens_seconds and llamacpp:predicted_tokens_seconds, as well as request and context counters. Compare prompt throughput with generation throughput to identify which phase changed. Record the runtime build, model and quantization, context length, concurrency, and relevant logs so you can reverse a change that hurts quality or performance.

Quick Recap

Bestseller No. 1
ASUS Dual Radeon RX 9060 XT 16GB GDDR6 Gaming Graphics Card
ASUS Dual Radeon RX 9060 XT 16GB GDDR6 Gaming Graphics Card
0dB technology lets you enjoy light gaming in relative silence; Dual BIOS switch lets you toggle between Quiet and Performance BIOS profiles
$529.00
Bestseller No. 2
GIGABYTE GeForce RTX 5070 Ti Gaming OC 16G Graphics Card, 16GB 256-bit GDDR7, PCIe 5.0, WINDFORCE Cooling System, GV-N507TGAMING OC-16GD Video Card
GIGABYTE GeForce RTX 5070 Ti Gaming OC 16G Graphics Card, 16GB 256-bit GDDR7, PCIe 5.0, WINDFORCE Cooling System, GV-N507TGAMING OC-16GD Video Card
Powered by the NVIDIA Blackwell architecture and DLSS 4; Powered by GeForce RTX 5070 Ti; Integrated with 16GB GDDR7 256bit memory interface
SaleBestseller No. 3
GIGABYTE GeForce RTX 5060 WINDFORCE OC 8G Graphics Card, Cooling System, 8GB 128-bit GDDR7, PCIe 5.0, Manufactured by NVIDIA, DisplayPort & HDMI - Video Output Interface, GV-N5060WF2OC-8GD Video Card
GIGABYTE GeForce RTX 5060 WINDFORCE OC 8G Graphics Card, Cooling System, 8GB 128-bit GDDR7, PCIe 5.0, Manufactured by NVIDIA, DisplayPort & HDMI - Video Output Interface, GV-N5060WF2OC-8GD Video Card
Powered by the NVIDIA Blackwell architecture and DLSS 4; Powered by GeForce RTX 5060; Integrated with 8GB GDDR7 128bit memory interface
$459.99
SaleBestseller No. 4
GIGABYTE Radeon RX 9070 XT Gaming OC 16G Graphics Card, PCIe 5.0, 16GB GDDR6, GV-R9070XTGAMING OC-16GD Video Card
GIGABYTE Radeon RX 9070 XT Gaming OC 16G Graphics Card, PCIe 5.0, 16GB GDDR6, GV-R9070XTGAMING OC-16GD Video Card
Powered by Radeon RX 9070 XT; WINDFORCE Cooling System; Hawk Fan; Server-grade Thermal Conductive Gel
$840.00
SaleBestseller No. 5
ASUS Prime Radeon RX 9070 XT 16GB GDDR6 OC Edition Gaming Graphics Card
ASUS Prime Radeon RX 9070 XT 16GB GDDR6 OC Edition Gaming Graphics Card
0dB technology lets you enjoy light gaming in relative silence; Dual BIOS switch lets you toggle between Quiet and Performance BIOS profiles
$829.00

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.