Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

You can run some AI models locally with a CPU and system RAM; a discrete GPU is not mandatory. The hardware you need depends on the model, its quantization, the context length, and how quickly you want it to respond. For GPU inference, plan for VRAM beyond the model weights; for CPU inference, the model uses system memory. Hybrid CPU/GPU inference can make larger models fit, usually with a performance trade-off.

Start with the model, not a universal RAM or VRAM minimum

There is no single memory threshold that guarantees every model will run. Choose the model and runtime first, then check the downloadable model’s weight size, quantization, supported hardware, and intended context length. Leave capacity for runtime buffers, the context’s key/value (KV) cache, the operating system, and any simultaneous requests. Longer contexts and parallel requests can increase memory use; see Ollama’s context-length guidance.

Quantization reduces the memory used for model weights, but can affect output quality. The trade-off depends on the model and task, so a smaller quantized version is not automatically the best choice. The llama.cpp project documents quantization options ranging from 1.5-bit to 8-bit; these options do not imply identical quality or compatibility across models.

How much memory does local AI need?

A model’s file size is only part of the budget. In a configuration-specific estimate in its gpt-oss guide, llama.cpp lists gpt-oss 20B at 12.0 GB of model data, 2.7 GB of compute buffers, and 0.2 GB of KV cache with an 8,192-token context: 14.9 GB total. At 131,072 tokens, the estimate rises to 17.9 GB.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
GMKtec EVO-X2 AI Mini PC Ryzen Al Max+ 395 Superchip 128GB LPDDR5X 2TB SSD
  • EVOLUTION RYZEN AI MAX+ 395 MINI PC - GMKtec EVO-X2 is the next evolution in AI mini PC Ryzen Strix Halo series. Thanks to AMD Simultaneous Multithreading (SMT) the core-count is effectively doubled, to 32 threads. Ryzen AI Max+ 395 has 64 MB of L3 cache and can boost up to 5.1 GHz, depending on the workload. The Ryzen AI Max+ 395 is currently rated as the "most powerful x86 APU" on the market for AI computing.
  • AI NPU with XDNA 2 ARCHITECTURE - Powered by 16 “Zen 5” CPU cores, 50+ peak AI TOPS XDNA 2 NPU and a truly massive integrated GPU driven by 40 AMD RDNA 3.5 CUs, the Ryzen AI MAX+ 395 is a transformative upgrade and delivers a significant performance boost over the competition. The Ryzen AI Max+ 395 excels in consumer AI workloads like the llama.cpp-powered application: LM Studio. Shaping up to be the must-have app for client LLM workloads, LM Studio allows users to locally run the latest language model without any technical knowledge required and unleash their creativity and productivity.
  • AMD RADEON 8090S iGPU GAMING PC - The AMD Radeon RX 8060S offers all 40 CUs with up to 2.9 GHz graphics clock and uses the new RDNA 3.5 architecture. The powerful iGPU is positioned between an RTX 4060 and 4070 laptop GPU and therefore enables gaming in FHD at maximum details in most demanding games. The 8060S can also utilize the full 128GB pool, which is perfect for running LLMs such as Deepseek 70B Q8, which runs comfortably on this machine.
  • EIGHT CHANNEL LPDDR5X - LPDDR5X is a new ground breaking memory small form factor installed on-board. With blazing speeds up to to 8000MT/s, it runs 1.5x faster than the DDR5 SODIMMs; 90% better performance over DDR5 SODIMMs in video conferencing and photo editing; 30% better performance in productivity apps; 12% better performance in digital content workloads.
  • QUAD SCREEN 8K DISPLAY SUPPORT - EVO-X2 AI Mini PC support 4-screen 4K/8K output via HDMI 2.1 (8K@60Hz), DisplayPort 1.4 (4K@60Hz), and dual USB 4 40Gbps Transfer speed (supporting PD3.0/DP1.4/DATA). Ideal for gaming, video editing, and multitasking, it provides expansive and crisp multi-display support.

For gpt-oss 120B, the same guide estimates 61.0 GB of model data, 2.7 GB of compute buffers, and 0.3 GB of KV cache at 8,192 tokens, or 64.0 GB total. At 131,072 tokens, it estimates 68.5 GB. These figures describe the guide’s configurations, not universal requirements; its CLI settings can change the estimates.

Runtime settings also matter. Ollama documents context defaults of 4k tokens below 24 GiB of VRAM, 32k for 24–48 GiB, and 256k at 48 GiB or more. These are Ollama defaults, not general hardware requirements or a promise that a particular model supports those context lengths.

Choose a hardware path

Hardware path What it offers What to check
CPU-only computer Runs compatible models without a discrete graphics card, using system memory. Capacity and speed depend on the CPU, available RAM, model, and runtime. No universal speed estimate applies.
Desktop with a discrete GPU A supported GPU backend can accelerate inference; VRAM determines how much of the model can stay on the GPU. Match VRAM, runtime/backend support, model, and context. Model weights are not the only memory demand.
Apple Silicon llama.cpp supports Apple Silicon, with ARM, Accelerate, and Metal optimizations. CPU and GPU share unified memory. Consider total unified memory and other system use; it is shared memory, not dedicated VRAM.
Hybrid CPU/GPU Partial GPU offload can run a model that does not fit entirely in VRAM. It can extend capacity, but performance depends on workload and configuration; do not assume it matches full GPU residency.
Intel or another supported accelerator llama.cpp lists Intel SYCL and OpenVINO support for Intel CPUs, GPUs, and NPUs, as well as Vulkan and other backends. Verify the exact device, driver, runtime, model format, and required features. Support in one backend does not establish support in another.

llama.cpp lists CUDA for NVIDIA, HIP for AMD, Metal for Apple Silicon, SYCL for Intel GPUs, and Vulkan among its backends. This shows that several hardware paths are available, not that every model works equally well on every device. Check the runtime’s current support details and the model’s requirements; llama.cpp’s documentation describes its backends and hybrid inference.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What to prioritize when buying or upgrading

  • For GPU inference: prioritize a graphics card with enough VRAM for the chosen model, context, buffers, and other GPU workloads, and confirm that your runtime supports its backend.
  • For CPU or hybrid inference: a desktop memory upgrade may help provide room for model data and runtime use. System RAM does not become dedicated GPU VRAM.
  • For model storage: an SSD can hold downloaded model files and help with storage capacity. It does not increase inference compute or replace the memory needed while a model runs.
  • For any path: leave headroom for the operating system and other workloads, and check what happens when the model does not fit entirely in GPU memory. CPU offload may enable execution but changes performance.

Ollama’s GPU documentation includes an NVIDIA GeForce RTX 4090 configuration example. It is an example of supported hardware, not a recommendation that it is the best choice for every workload or budget.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Practical way to size your system

  1. Pick the model and runtime. Confirm the model format and quantization you intend to use, and check the runtime’s supported backends and device requirements.
  2. Set the context and workload. Decide how much context you need and whether requests will run concurrently; both can add memory demand.
  3. Estimate total memory, not just the download. Account for weights, runtime buffers, KV cache, operating-system use, and other active work. Treat published examples as configuration-specific.
  4. Match memory to the execution path. For GPU use, compare the estimate with VRAM; for CPU use, compare it with available system RAM. For Apple Silicon, account for shared unified memory.
  5. Check compatibility and fallback behavior. Confirm the exact runtime, accelerator, driver, and model format. If the model exceeds VRAM, determine whether CPU offload is supported and whether its likely trade-off suits your needs.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.