Quick wins for a faster PC:
Repair Windows errors before they cause bigger problemsFix Now →Scan for outdated or missing drivers - takes under a minuteDriver Scan →Clear out junk files and repair common Windows errorsFree Scan →Find out which part of local inference is slow before changing settings: model loading, prompt processing, or token generation. Then check the runtime’s logs and memory use to see whether the cause is GPU placement, context and KV-cache allocation, CPU threading, or a model that exceeds available memory. Ollama, llama.cpp, and vLLM use different controls, so verify commands against the version you have installed.
Identify which part of inference is slow
“Slow inference” can describe three different delays. Separate them before tuning, because a fix for model loading will not necessarily improve token generation.
- Model loading: the delay before the runtime says the model is ready.
- Prompt processing: the delay before the first generated token, while the model reads the input.
- Token generation: the rate at which new tokens appear after generation starts.
If model loading is slow
For vLLM, model downloads depend on network conditions. Large model files, shared or network filesystems, and host-memory pressure can also slow loading; swapping to disk can make it especially sluggish. Check whether the model is already local, whether its files are on a local disk, and whether system memory is under pressure. The vLLM v0.18.2 troubleshooting guide suggests using a local model path and local disk, monitoring CPU memory, and trying --load-format dummy to isolate loading behavior.
If the first token is slow but later tokens are not
Look at prompt processing and context size. A long prompt requires more processing before generation begins, and the selected context length also affects memory allocation. If the runtime exposes separate prompt and generation metrics, compare them rather than treating both as one speed number.
#1 Best Overall
- Axial-tech fans now feature a smaller fan hub that facilitates longer blades and a barrier ring that increases downward air pressure
- 2.5-slot design allows for greater build compatibility while maintaining cooling performance
- 0dB technology lets you enjoy light gaming in relative silence
- Dual BIOS switch lets you toggle between Quiet and Performance BIOS profiles
- Dual ball fan bearings last up to twice as long as sleeve bearing designs
If every generated token is slow
Check whether the model is actually running on the GPU, whether some layers remain on the CPU, and whether CPU threads or temporary debugging instrumentation are limiting performance. The sections below explain how to verify those conditions.
Check whether the GPU is doing inference work
Seeing a GPU in the system does not prove the runtime is using it. Confirm actual placement in startup logs. In CUDA-enabled llama.cpp runs, look for diagnostics showing how many layers were offloaded and how much VRAM they use. The llama.cpp performance guide identifies these startup messages as evidence of GPU use.
In llama.cpp, -ngl (also available as --gpu-layers) requests GPU offload. A large value asks the runtime to offload as many layers as fit; it does not guarantee that every layer will fit. If logs show partial offload, CPU-resident layers may constrain speed. Confirm that your installed build supports the device backend, then check the startup output rather than assuming a flag succeeded.
Rank #2
- Powered by the NVIDIA Blackwell architecture and DLSS 4
- Powered by GeForce RTX 5070 Ti
- Integrated with 16GB GDDR7 256bit memory interface
- PCIe 5.0
- WINDFORCE cooling system
The llama.cpp server CLI reference documents --fit, enabled by default, as a setting that adjusts unspecified arguments to fit device memory. It also describes multi-GPU placement options: layer split (the documented default), row split, and experimental tensor split. These modes distribute work differently; use the current CLI reference for your build before changing them. Flags for llama.cpp should not be copied to Ollama or vLLM.
Recommended Free Tools
Reduce context and KV-cache memory carefully
Context length is not just a prompt-capacity setting: the KV cache used to retain context consumes memory. If you set a much longer context than your task needs, reduce it to a realistic length first. This can relieve memory pressure, though it also limits how much conversation or source text the model can retain.
Ollama Flash Attention
Ollama documents Flash Attention as a way to significantly reduce memory use as context grows when the selected backend and devices support it. To force it, set OLLAMA_FLASH_ATTENTION=1; set it to 0 to disable it. See the Ollama FAQ for the documented conditions and settings.
Rank #3
- Powered by the NVIDIA Blackwell architecture and DLSS 4
- Powered by GeForce RTX 5060
- Integrated with 8GB GDDR7 128bit memory interface
- PCIe 5.0
- WINDFORCE cooling system
Ollama KV-cache precision
With Flash Attention enabled, Ollama documents the OLLAMA_KV_CACHE_TYPE setting. Its FAQ describes these approximate memory and quality tradeoffs:
| Ollama cache type | Documented memory use relative to f16 | Documented quality tradeoff |
|---|---|---|
f16 |
Default; reference value | Reference precision |
q8_0 |
Approximately half of f16 | Very small loss |
q4_0 |
Approximately one quarter of f16 | Small-to-medium loss, potentially more noticeable at higher context sizes |
These are Ollama’s approximate, project-documented comparisons, not guaranteed measurements for every model or runtime. Ollama says the response-quality effect depends on the model and task; models with a high GQA count may be more sensitive to reduced precision. Try a cache setting on representative prompts and check output quality rather than assuming the smallest cache is best.
Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallDiagnose out-of-memory errors
An OOM error means a runtime allocation could not be satisfied, but the message alone does not identify what consumed the memory. Check runtime logs and resource use to distinguish among model weights, KV cache, concurrency, and other runtime allocations. Also identify whether the shortage is in GPU memory or host RAM: the remedies differ.
Rank #4
- Powered by Radeon RX 9070 XT
- WINDFORCE Cooling System
- Hawk Fan
- Server-grade Thermal Conductive Gel
- RGB Lighting
vLLM states that “If the model is too large to fit in a single GPU, you will get an out-of-memory (OOM) error.” Its troubleshooting documentation points users toward memory-reduction options. There is no universal VRAM threshold for local LLMs: fit depends on the model, its representation, context and runtime allocation, and the runtime’s placement options.
Try lower-impact changes first
- Reduce unnecessary context and concurrency. Use only the context your task requires, and reduce simultaneous requests if the runtime’s memory use rises with concurrent work.
- Reduce KV-cache use where supported. For Ollama, consider the documented Flash Attention and cache-precision settings, then validate response quality.
- Choose a smaller or lower-memory model representation. Use a model or quantization supported by your chosen runtime. This changes the model representation and may affect output quality; compare results on the work you actually do.
- Adjust supported GPU placement or splitting. For llama.cpp, inspect GPU offload and consult the installed build’s options, including its fit and multi-GPU controls. Do not assume another runtime accepts the same flags.
- Add hardware capacity only after identifying the constrained resource. More GPU memory may help when weights or cache do not fit, but it is not a diagnosis or a universal fix; first establish the model, context, runtime, and workload involved.
Tune CPU threads without oversaturating the processor
If llama.cpp is doing CPU work, or CPU threads are part of a mixed CPU/GPU run, too many threads can oversaturate the processor and reduce performance. The llama.cpp performance guide suggests starting with one thread and doubling the count until a bottleneck appears, then scaling back. If one thread helps, it also suggests explicitly trying the number of physical CPU cores. Treat this as a troubleshooting heuristic, not a universally optimal setting.
The same guide reports a project benchmark of 9.1 tokens per second for a specific setup: an NVIDIA A6000 with 48 GB VRAM, a CPU with seven physical cores, 32 GB RAM, and a 30B-parameter Q4_0 GGML model. In its listed results, -t 4 with the stated large GPU-layer setting measured 9.1 tokens per second, while -t 7 with that setting measured 8.7 tokens per second. The page does not state a publication year. These figures describe that configuration, not expected performance on a different machine or with a current model format.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Best Value
- Axial-tech fans now feature a smaller fan hub that facilitates longer blades and a barrier ring that increases downward air pressure
- Phase-change GPU thermal pad helps ensure optimal heat transfer, lowering GPU temperatures for enhanced performance and reliability
- 2.5-slot design allows for greater build compatibility while maintaining cooling performance
- Dual-ball fan bearings last up to twice as long as standard conventional sleeve bearings designs
- 0dB technology lets you enjoy light gaming in relative silence
Turn off temporary debugging instrumentation
Debug settings can distort performance after the original issue has been diagnosed. vLLM warns that VLLM_TRACE_FUNCTION=1 slows token generation by over 100x and should not be used unless absolutely needed. Turn off that setting and other debugging environment variables that you enabled for troubleshooting, then measure again. See the vLLM troubleshooting guide.
Measure the change, not just the impression
Use the same model, prompt, context, and workload when comparing settings; otherwise, a faster result may simply reflect less work. The llama.cpp server reference exposes separate metrics, including llamacpp:prompt_tokens_seconds and llamacpp:predicted_tokens_seconds, as well as request and context counters. Compare prompt throughput with generation throughput to identify which phase changed. Record the runtime build, model and quantization, context length, concurrency, and relevant logs so you can reverse a change that hurts quality or performance.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

