Do these 3 things before closing this tab:
1Scan for outdated or missing drivers - takes under a minute2Repair Windows errors before they cause bigger problems3Fix the driver behind crashes, sound loss and screen glitchesChoose a GPU by checking whether its usable memory can hold your exact model, precision, context length and expected workload—with room for inference overhead. Only after that should you compare speed, software support, power and price. Parameter count alone cannot tell you whether a model will fit or run well.
This guide focuses on inference: loading a model to generate responses, rather than training or fine-tuning it. The memory estimates are starting points, not guarantees; the model architecture and inference software affect the result.
How much VRAM do you need to run an AI model?
Start with the memory occupied by the model’s weights. Hugging Face’s Transformers optimization documentation gives a rule of thumb: a model with X billion parameters needs about 2 × X GB for weights in bfloat16 or float16, or about 4 × X GB in float32. Those figures cover weights only, not every allocation needed to run inference. Hugging Face’s documentation explains the estimate and its limits.
| Parameter count | Approx. bfloat16/float16 weight memory | Approx. float32 weight memory |
|---|---|---|
| 7 billion | 14 GB (about 2 GB per billion parameters; Hugging Face rule of thumb) | 28 GB (about 4 GB per billion parameters; Hugging Face rule of thumb) |
| 13 billion | 26 GB (about 2 GB per billion parameters; Hugging Face rule of thumb) | 52 GB (about 4 GB per billion parameters; Hugging Face rule of thumb) |
| 70 billion | 140 GB (about 2 GB per billion parameters; Hugging Face rule of thumb) | 280 GB (about 4 GB per billion parameters; Hugging Face rule of thumb) |
The table applies the documentation’s approximate formula; it is not a list of measured GPU requirements. In practice, memory is usually specified in GB on product pages, and usable capacity can be lower than the card’s nominal total because the runtime, display, or other GPU work also needs memory.
Recommended Free Tools
#1 Best Overall
- Powered by the NVIDIA Blackwell architecture and DLSS 4
- Powered by GeForce RTX 5080
- Integrated with 16GB GDDR7 256bit memory interface
- PCIe 5.0
- WINDFORCE cooling system
Budget for context and concurrent requests
Model weights are only part of the inference footprint. A longer prompt or generated sequence can require more attention memory, including memory for the key-value (KV) cache. Serving multiple requests at once and runtime allocations can add to demand as well. The amount varies by model architecture and software, so there is no reliable universal percentage to add to the weight estimate. Hugging Face discusses how sequence length affects attention memory and the KV cache in its optimization guide.
When estimating fit, use the context length you intend to run—not only the model’s advertised maximum—and account for how many requests may overlap. A configuration that works for a short, single-user prompt may not fit the same model at a longer context or with concurrent requests.
Rank #2
- Powered by Radeon RX 9070 XT
- WINDFORCE Cooling System
- Hawk Fan
- Server-grade Thermal Conductive Gel
- RGB Lighting
Count all parameters for a mixture-of-experts model
For a mixture-of-experts (MoE) model, only some experts may be active for a given token, but that does not mean the GPU only needs memory for those active parameters. Experts must be available to the model even when a particular token uses only a subset. NVIDIA’s technical explanation distinguishes active parameters from the total model and discusses deployment considerations for dense and MoE models: Dense vs. MoE Models.
Can your GPU run this model?
Use the exact checkpoint and runtime you plan to use. A model name or parameter count by itself is not enough: weight format, quantization, architecture, context length, request concurrency and software all affect whether it fits.
Rank #3
- AMD Radeon RX 550 Chipset, Silver plated PCB & all solid capacitors provide lower temperature, higher efficiency & stability
- 9CM unique fan provide low noise and huge airflow for your GPU
- GPU Boost Clock / Memory Speed : up to 1183 MHz / 4GB GDDR5 / 6000 MHz Memory, Stream Processors 512, Perfect for 3D CAD/CAM working, video and photo editing, Video Games @1080p
- Support: DirectX 12, Shader Model 5.0, OpenGL 4.6/4.5, 4K Video Decode
- Identify the checkpoint and architecture. Record the model variant and parameter count. For MoE models, do not use active parameters per token as a substitute for the total model when estimating memory.
- Choose a weight format or quantization. Estimate weight memory at that precision, using the model’s documented requirements where available. Quantization can reduce memory substantially, but its effect varies by checkpoint and method.
- Set the intended workload. Note the context length, expected simultaneous requests and any other GPU tasks. Include image or video resolution if the model handles visual inputs.
- Compare the total need with usable GPU memory. Leave capacity for the KV cache, runtime allocations and other GPU use. If the estimate is close to the card’s limit, test the actual setup rather than assuming it will fit.
- Confirm the software path. Check that the operating system, GPU architecture, drivers, inference backend and model format work together. NVIDIA’s local-AI guidance likewise recommends defining target VRAM and performance needs and selecting a backend based on the system and workload: Build Local AI With NVIDIA GPUs.
Can quantization make a smaller GPU work?
Often, yes: quantization stores weights at lower precision to reduce memory demand. But it does not guarantee a particular model will fit, nor does a given bit depth imply the same quality or speed across different quantizers and checkpoints. Evaluate the exact model on the task you care about.
Hugging Face’s documented OctoCoder example uses about 32 GB in its baseline, 15 GB at 8-bit, and a little over 9 GB at 4-bit. These are figures for that example, not general requirements for other models. The same documentation warns that quantization trades memory efficiency against accuracy and, in some cases, inference time. See the OctoCoder example and quantization discussion.
Rank #4
- Powered by the NVIDIA Blackwell architecture and DLSS 4. System Requirements: Minimum 850W PSU with 16-pin 12V-2x6 (12VHPWR) connector required. Verify before purchasing.
- Military-grade components deliver rock-solid power and longer lifespan for ultimate durability. Compatibility: 348mm (13.7") length, 3.6 slots, 4.3 lbs. Confirm case clearance and slot spacing. GPU bracket included.
- Protective PCB coating helps protect against short circuits caused by moisture, dust, or debris
- 3.6-slot design with massive fin array optimized for airflow from three Axial-tech fans
- Phase-change GPU thermal pad helps ensure optimal thermal performance and longevity, outlasting traditional thermal paste for graphics cards under heavy loads
Before choosing a card on the assumption that a quantized version will be usable, confirm that the precise checkpoint and quantization format are supported by your intended backend. Then assess output quality and speed on representative prompts; fitting in memory is necessary, but does not establish that the result meets your needs.
What GPU should you buy to run local AI models?
There is no single best GPU for every open-weights model. First choose a memory capacity that fits your target workload with headroom; then compare measured speed and compatibility for that model and software. Once those are satisfactory, consider power, cooling, physical fit, platform requirements, price and local availability.
Free tools Windows power users keep installed
One-click scans. No signup required.
Best Value
- System Compatibility Note: This 2‑slot card measures 249 mm (L) x 132 mm (W) x 41 mm (H) and requires a single 8‑pin power connector. Please verify available chassis clearance and ensure your power supply is rated for a recommended 550W before purchase.
- Dedicated Support: Please contact us directly through Amazon for any product questions or assistance you may require.
- Next‑Gen AMD RDNA 4 Architecture: Powered by the AMD Radeon RX 9060 XT GPU with 32 Compute Units featuring 3rd Gen Ray Tracing and 2nd Gen AI Accelerators, delivering exceptional 1440p gaming and AI‑enhanced performance.
- Blazing‑Fast Engine Clock: Delivers a boost clock of up to 3290 MHz and a game clock of 2700 MHz out of the box, providing the raw power for smooth, high‑framerate gameplay.
- 16GB GDDR6 Memory on 128‑Bit Bus: Equipped with 16GB of high‑speed GDDR6 memory running at 20 Gbps, offering ample capacity and bandwidth for modern game textures and creative applications.
- Usable VRAM: Match capacity to the model, selected precision, context and concurrency. Treat a close fit as a warning to test, not as a safe margin.
- Performance: Compare memory bandwidth and workload-specific latency or throughput, such as tokens per second, using results for the model and settings you expect to run. Benchmark claims are meaningful only alongside details such as model, quantization, software, driver, prompt and system configuration.
- Compatibility: Verify the inference backend supports the GPU architecture, model format and operating system. A card’s theoretical capacity does not ensure that your chosen runtime can use it effectively.
- System constraints: Check power delivery, cooling, card dimensions and the rest of the platform. A card that fits the memory target may still be unsuitable for the computer it must run in.
- Cost and availability: Compare current local listings only after confirming the above. The cited sources establish no current market-wide price comparison or GPU ranking.
Vendor documentation can help verify specifications and reproduce a stated test, but results from different vendors or systems are not directly comparable without matching configurations. AMD, for example, identifies the Radeon AI PRO R9700 as a 32 GB card and documents local-inference tests with named quantized models and system and software details in its Radeon AI PRO ROCm PyTorch guide. That documentation is an example of a card and test configuration, not evidence that it is the best-value or best-performing choice for every workload.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Will multiple GPUs or shared system memory solve a capacity limit?
Multiple GPUs
Model-parallel approaches can divide a model across more than one GPU when it will not fit on a single card. They add setup and communication considerations, however, and do not automatically make inference faster. Hugging Face notes that naïvely placing layers across GPUs can leave some devices idle. Check how the selected backend distributes the model and measure the full workload, including any communication overhead, rather than assuming that adding cards scales performance linearly. Hugging Face’s optimization documentation covers model placement and memory considerations.
Unified-memory systems
Some systems let integrated graphics use a portion of system RAM as graphics memory. AMD’s Ryzen AI Max+ example describes up to 96 GB of Variable Graphics Memory on a 128 GB Ryzen AI Max+ 395 platform. AMD also says memory assigned to VGM is no longer available as CPU system RAM. This expands the allocation available to graphics, but should not be treated as equivalent in speed or behavior to discrete GPU VRAM without evidence for the particular workload. AMD’s VGM FAQ describes the platform and tradeoff.
A practical way to make the final choice
Write down the configuration you actually want to use, then rule out cards that cannot meet it before comparing speed or price:
- Model checkpoint, architecture and total parameter count.
- Weight format or quantization and the backend that supports it.
- Target context length, request concurrency and any visual-input resolution.
- Estimated weight memory plus workload-specific room for KV cache and runtime allocations.
- Measured performance for a comparable model, precision, prompt and system configuration.
- Operating-system and driver support, power and cooling requirements, physical dimensions, and current local cost.
If no single card provides enough usable memory, compare a supported multi-GPU setup with a system that can reallocate shared memory. Treat each as a distinct option with its own software, performance and system-memory tradeoffs. Re-test the final configuration with the exact checkpoint and workload before relying on it.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

