Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

If Ollama is slow or using more memory than expected, first check what it actually loaded: run ollama ps while the affected model is running and inspect its processor split and context size. That tells you whether the immediate issue is CPU offload, a large context, or memory pressure from parallel requests—not simply whether a GPU setting is enabled.

Start with what Ollama actually loaded

Reproduce the slowdown or high-memory use, then open a terminal and run:

ollama ps

Check the PROCESSOR and CONTEXT columns. The processor value indicates whether the model is running on the GPU, CPU, or a mix; the context value shows the context allocated for that running model. Record the model, these values, whether other models are loaded, and whether other requests are active. Ollama’s context-length documentation advises avoiding CPU offload for best performance where possible.

A GPU being present—or an app setting that appears to enable GPU use—does not establish where this particular model is running. If the processor split is unexpected, follow the GPU-discovery checks below. If the model is on the GPU but memory is tight, start by reducing unnecessary context and concurrency.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Ollama is very slow: check context and CPU offload

Context is the amount of token history the model can access in memory. Ollama describes it as “the maximum number of tokens that the model has access to in memory.” Larger context lengths require more memory, and can make an otherwise workable model exceed available GPU memory and spill work onto the CPU.

Ollama’s current documentation lists these default context lengths by available VRAM tier. These are documented defaults, not a guarantee that every model will fit or a recommendation for every task; defaults can change.

Available VRAM Documented default context length
Below 24 GiB 4k
24–48 GiB 32k
48 GiB or more 256k

The context guide recommends at least 64,000 tokens for tasks such as agents, web search, and coding tools. That larger context has a corresponding memory cost. Keep the context large enough for the work, but do not allocate more than the task needs.

Ways to lower context

Depending on how you run Ollama, change the context using the Ollama app setting, the OLLAMA_CONTEXT_LENGTH environment variable, or a runtime parameter. Use the setting appropriate to your installation and check ollama ps again after reloading the model to verify the allocated value changed.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

If the processor split shows CPU use that does not match your system, do not assume context is the only cause. Check the server logs and platform-specific GPU discovery steps in the section below.

Ollama uses too much memory: reduce parallel requests

Memory use can rise when Ollama serves several requests at once. Its FAQ says required RAM scales with OLLAMA_NUM_PARALLEL * OLLAMA_CONTEXT_LENGTH: parallel processing increases allocated context in proportion to the number of simultaneous requests.

When memory is constrained, lower OLLAMA_NUM_PARALLEL or avoid keeping multiple models loaded at the same time. Fewer parallel requests can reduce memory pressure, but also reduce throughput when multiple users or tasks need responses simultaneously. Defaults and controls can vary by deployment, so check your installed Ollama version and configuration rather than assuming a particular default.

Consider Flash Attention and KV-cache options

Ollama says Flash Attention can significantly reduce memory use as context grows. When the selected backend and devices support it, Ollama uses Flash Attention automatically. Its benefits therefore depend on the model and hardware/backend in use.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

With Flash Attention enabled, Ollama documents three KV-cache types. These ratios apply to KV-cache memory, not total model memory.

Rank #4
Sale
HPE NVIDIA Tesla V100 32GB HBM2 PCIe 3.0 x16 Passive GPU Computational Accelerator for AI Machine Learning HPC Deep Learning 699-2G500-0216-400 (Renewed)
  • NVIDIA Volta GV100 Architecture — 4,608 CUDA Cores, 640 1st-Gen Tensor Cores delivering 14 TFLOPS FP32 and 112 TFLOPS deep learning performance for AI training, inference, HPC, and scientific computing workloads
  • 32GB HBM2 ECC Memory — 900 GB/s Bandwidth — High-bandwidth memory on a 4096-bit bus with ECC error correction provides the memory capacity and throughput required for the largest AI models, simulations, and datasets
  • PCIe 3.0 x16 Interface — 250W TDP — Standard PCIe Gen3 connectivity with passive cooling designed for enterprise rack server deployment in HPE ProLiant, Dell PowerEdge, and Supermicro platforms with adequate chassis airflow
  • NVLink — Scale to 96GB Unified Memory — Connect two V100 GPUs via NVLink at 300 GB/s bi-directional bandwidth to scale GPU memory from 32GB to 96GB for larger AI training and HPC workloads
  • Multi-Precision Computing — Supports FP64 (7 TFLOPS), FP32 (14 TFLOPS), FP16 (112 TFLOPS) and INT8 precision modes for flexible deployment across training, inference, and scientific simulation workloads
KV-cache type Memory use compared with f16 Documented quality tradeoff
f16 Default Baseline option
q8_0 Approximately half Usually no noticeable quality impact
q4_0 Approximately one quarter Small-to-medium quality loss, potentially more noticeable at high context

These are approximate FAQ figures; actual impact varies by model and task. KV-cache quantization is configured globally with OLLAMA_KV_CACHE_TYPE in the documented setup. Choose a lower-memory option only if its quality tradeoff is acceptable for your work.

GPU not being used or GPU not detected: check logs and platform

A model that does not fit in available GPU memory is a different problem from Ollama being unable to discover the GPU. If ollama ps shows unexpected CPU use, check the Ollama server log for your operating system and installation method, then match the troubleshooting steps to your hardware.

  • macOS: use the documented Ollama log location and confirm that the GPU/backend path applies to your Mac.
  • Linux: inspect the systemd journal when Ollama runs as a system service.
  • Docker: inspect the container logs and verify that the container runtime is configured to expose the GPU.
  • Windows: check Ollama’s documented log files.

Ollama’s troubleshooting guide also covers debug logging and vendor-specific discovery checks. For NVIDIA, relevant causes can include container runtime configuration, driver problems, or UVM. For AMD, check driver compatibility and device permissions, including access to /dev/kfd where applicable. Do not run privileged driver commands without first confirming they match your platform and the problem shown in the logs.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For compatibility, consult Ollama’s GPU support page for the current requirements and supported paths for your operating system and GPU generation. It documents NVIDIA compute-capability and driver requirements, Metal for Apple GPUs, and Vulkan support paths; the matrix may change.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

When a hardware upgrade may help

Consider hardware only after reducing context or concurrency that you do not need and confirming Ollama can see the GPU. More GPU memory may help if your actual workload still exceeds what the system can hold, but there is no universal model-to-VRAM fit or speed forecast: model size, quantization, context, concurrent work, competing GPU use, and the GPU/backend all matter.

Ollama’s support list includes the NVIDIA GeForce RTX 5060, but support does not establish that a particular card will fit your model at your required context or run at a guaranteed speed. Compare available VRAM and workload fit, then check Ollama compatibility, power delivery, case clearance, and cost for the exact system before buying. Ollama announced a new model scheduling system on September 23, 2025, describing more exact memory measurement and improvements to out-of-memory crashes, GPU allocation/utilization, and multi-GPU scheduling. Those benefits apply to models implemented in that engine; they should not be assumed for every model.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.