You can run an AI model on your own computer by installing a local runtime, downloading compatible model weights, and loading them in that runtime. For the easiest first run, use LM Studio’s graphical interface; choose Ollama for a simple command-line workflow, or llama.cpp if you want more control or a local server. Before downloading, check your computer’s memory and free disk space, and check the exact model’s format and license: “open-source” is often used loosely for models with downloadable weights, but access to weights does not mean every model has the same license or degree of openness.
What you need to run a model locally
A local runtime and a model are separate components. The runtime loads and runs the model; the model weights are the files containing the trained model. Install a runtime, then select and download weights it can use. Common formats include GGUF and SafeTensors, but format compatibility depends on the runtime and the specific model.
- A supported computer: Operating-system and processor requirements vary by runtime.
- Enough memory: RAM, GPU memory (VRAM), model quantization, and context length all affect whether a model loads and how it performs.
- Enough storage: Model downloads can be large. Ollama’s Windows documentation says downloaded models may take tens to hundreds of GB, depending on what you download.
- A model suited to your task: Check the exact variant for capabilities such as coding, image input, audio, long context, or tool use; a model family name alone does not guarantee them.
For a first attempt on an ordinary laptop, start with a smaller instruction-tuned model and a modest context setting. Move to a larger model only if your task needs it and your available memory can support it.
Run a model with LM Studio
LM Studio is the least technical option in this guide: you can find and download models in its interface, then load one for a chat. Check its current system requirements before installing. Its requirements are recommendations, not a guarantee that every model will fit or run smoothly.
#1 Best Overall
- EVOLUTION AMD RYZEN AI MAX+ 395 MINI PC - GMKtec EVO-X2 is the next evolution in AI mini PC Ryzen Strix Halo series. Thanks to AMD Simultaneous Multithreading (SMT) the core-count is effectively doubled, to 32 threads. Ryzen AI Max+ 395 has 64 MB of L3 cache and can boost up to 5.1 GHz, depending on the workload. The Ryzen AI Max+ 395 is currently rated as the "most powerful x86 APU" on the market for AI computing.
- AI NPU with XDNA 2 ARCHITECTURE - Powered by 16 “Zen 5” CPU cores, 50+ peak AI TOPS XDNA 2 NPU and a truly massive integrated GPU driven by 40 AMD RDNA 3.5 CUs, the Ryzen AI MAX+ 395 is a transformative upgrade and delivers a significant performance boost over the competition. The Ryzen AI Max+ 395 excels in consumer AI workloads like the llama.cpp-powered application: LM Studio. Shaping up to be the must-have app for client LLM workloads, LM Studio allows users to locally run the latest language model without any technical knowledge required and unleash their creativity and productivity.
- AMD RADEON 8090S iGPU GAMING PC - The AMD Radeon RX 8060S offers all 40 CUs with up to 2.9 GHz graphics clock and uses the new RDNA 3.5 architecture. The powerful iGPU is positioned between an RTX 4060 and 4070 laptop GPU and therefore enables gaming in FHD at maximum details in most demanding games. The 8060S can also utilize the full 128GB pool, which is perfect for running LLMs such as Deepseek 70B Q8, which runs comfortably on this machine.
- EIGHT CHANNEL LPDDR5X - LPDDR5X is a new ground breaking memory small form factor installed on-board. With blazing speeds up to to 8000MT/s, it runs 1.5x faster than the DDR5 SODIMMs; 90% better performance over DDR5 SODIMMs in video conferencing and photo editing; 30% better performance in productivity apps; 12% better performance in digital content workloads.
- QUAD SCREEN 8K DISPLAY SUPPORT - EVO-X2 AI Mini PC support 4-screen 4K/8K output via HDMI 2.1 (8K@60Hz), DisplayPort 1.4 (4K@60Hz), and dual USB 4 40Gbps Transfer speed (supporting PD3.0/DP1.4/DATA). Ideal for gaming, video editing, and multitasking, it provides expansive and crisp multi-display support.
- Check compatibility. LM Studio’s requirements say Apple Silicon Macs need macOS 14 or newer and recommend at least 16 GB of RAM. Macs with 8 GB may work with smaller models and modest context. On Windows, LM Studio supports x64 and Snapdragon X Elite ARM systems; x64 requires AVX2. It recommends 16 GB of RAM and at least 4 GB of dedicated VRAM.
- Install LM Studio. Download and install the version for your operating system from the LM Studio website.
- Find and download weights. Open Discover, search for a model, and choose a compatible variant. Confirm the format and the model’s license and terms before relying on it for a particular use.
- Load the model. Open Chat and select the downloaded model in the loader. Loading uses memory for the weights and other settings, including context.
- Start a conversation. If the model will not load, reduce the context setting or choose a smaller or more heavily quantized variant.
Run a model with Ollama from the command line
Ollama provides a command-line interface and a model library. Install it using the official download instructions for your operating system, then choose a model and variant from the Ollama library. The catalog changes, so select the current model name and variant there rather than relying on an old command copied from a guide.
- Install Ollama for your operating system using its official installer.
- Select a model from Ollama’s library, checking the current name, variant, capabilities, and originating model’s terms.
- Run the model using the command shown for that model in the current library entry.
- Check disk use before downloading. On Windows, Ollama runs as a native application and documents a local API at
http://localhost:11434. Its Windows documentation says the application binary needs at least 4 GB of storage, while downloaded models may need tens to hundreds of GB.
If you are short on internal storage, Ollama documents changing the model directory with the OLLAMA_MODELS environment variable. Check the exact model size and available space before downloading or relocating files. A catalog listing does not establish that every model is fully open source; check the originating model’s license and terms, especially for commercial use.
Rank #2
- Built for Local AI Development: AMD Ryzen AI Halo is designed for local AI development and inference, featuring 128GB unified memory and support for up to 200B parameter models to build and run intensive AI workloads locally.
- 128GB Unified Memory: Features 128GB LPDDR5x unified memory at 8000 MT/s with 256 GB/s memory bandwidth, providing a shared memory pool across the CPU, GPU, and NPU to support larger AI models.
- AMD Ryzen AI Max+ 395 Processor: Features 16 cores, 32 threads, and Zen 5 architecture, paired with AMD Radeon 8060S integrated graphics featuring 40 RDNA 3.5 compute units and an AMD XDNA 2 NPU with up to 50 TOPS.
- Linux AI Developer Platform: Purpose-built for Linux-based AI development with full AMD ROCm software support and preloaded tools, models, and workflows optimized for local AI development.
- Compact, Connected Design: Includes a 2TB M.2 SSD, 10GbE LAN, Wi-Fi 7, Bluetooth 5.4, USB-C connectivity, and HDMI 2.1b.
Use llama.cpp for command-line control or a local server
llama.cpp is a flexible runtime for GGUF models. Its project documents installation through package managers, Docker, prebuilt releases, or a source build, and supports CPU and multiple accelerator backends as well as hybrid CPU/GPU inference. Follow the llama.cpp project instructions for your platform and confirm the current syntax and supported backend before running commands.
For a GGUF file already on your computer, the project README gives this local-file example:
Recommended Free Tools
Rank #3
- EVOLUTION RYZEN AI MAX+ 395 MINI PC - GMKtec EVO-X2 is the next evolution in AI mini PC Ryzen Strix Halo series. Thanks to AMD Simultaneous Multithreading (SMT) the core-count is effectively doubled, to 32 threads. Ryzen AI Max+ 395 has 64 MB of L3 cache and can boost up to 5.1 GHz, depending on the workload. The Ryzen AI Max+ 395 is currently rated as the "most powerful x86 APU" on the market for AI computing.
- AI NPU with XDNA 2 ARCHITECTURE - Powered by 16 “Zen 5” CPU cores, 50+ peak AI TOPS XDNA 2 NPU and a truly massive integrated GPU driven by 40 AMD RDNA 3.5 CUs, the Ryzen AI MAX+ 395 is a transformative upgrade and delivers a significant performance boost over the competition. The Ryzen AI Max+ 395 excels in consumer AI workloads like the llama.cpp-powered application: LM Studio. Shaping up to be the must-have app for client LLM workloads, LM Studio allows users to locally run the latest language model without any technical knowledge required and unleash their creativity and productivity.
- AMD RADEON 8090S iGPU GAMING PC - The AMD Radeon RX 8060S offers all 40 CUs with up to 2.9 GHz graphics clock and uses the new RDNA 3.5 architecture. The powerful iGPU is positioned between an RTX 4060 and 4070 laptop GPU and therefore enables gaming in FHD at maximum details in most demanding games. The 8060S can also utilize the full 128GB pool, which is perfect for running LLMs such as Deepseek 70B Q8, which runs comfortably on this machine.
- EIGHT CHANNEL LPDDR5X - LPDDR5X is a new ground breaking memory small form factor installed on-board. With blazing speeds up to to 8000MT/s, it runs 1.5x faster than the DDR5 SODIMMs; 90% better performance over DDR5 SODIMMs in video conferencing and photo editing; 30% better performance in productivity apps; 12% better performance in digital content workloads.
- QUAD SCREEN 8K DISPLAY SUPPORT - EVO-X2 AI Mini PC support 4-screen 4K/8K output via HDMI 2.1 (8K@60Hz), DisplayPort 1.4 (4K@60Hz), and dual USB 4 40Gbps Transfer speed (supporting PD3.0/DP1.4/DATA). Ideal for gaming, video editing, and multitasking, it provides expansive and crisp multi-display support.
llama-cli -m my_model.gguf
To start a server using its documented Hugging Face syntax, the README gives:
llama-server -hf ggml-org/gemma-3-1b-it-GGUF
Use a model file and backend compatible with your build. CPU/GPU hybrid inference means that compatibility with an accelerator does not necessarily mean every model layer runs on the GPU. A server may make a model available to other applications, so understand its access controls before exposing it beyond your computer.
Rank #4
Choose a model and estimate memory
Model parameter count alone is not enough to predict whether a model will run well. Compare the exact variant against your machine and task using these factors:
- Runtime and format: Make sure the chosen runtime supports the model’s file format and variant.
- Memory and quantization: Quantized weights can reduce memory use, often with trade-offs in output quality. Leave room for runtime overhead and context, not just the weights.
- Context length: Longer context uses additional memory. A large advertised context window does not mean your computer can use it at full length.
- Task capability: Verify the specific variant supports the coding, image, audio, or tool/function features you need.
- Actual speed: Performance depends on the computer, runtime, backend, and settings. There is no universal performance winner among these runtimes.
- License and terms: Read the exact model card and license, particularly before commercial use.
Google’s Gemma 4 documentation illustrates how precision changes loading memory. The following are Google’s approximate GPU/TPU loading estimates, published in 2026 and last updated 2026-07-08 UTC. They include the page’s stated 20% overhead for additional loading items, but exclude supporting software and context-window memory. Google notes actual requirements vary by inference tool and environment, and longer context increases memory use. These figures apply to the listed Gemma 4 variants only; they are not a general calculator for other models.
Free tools Windows power users keep installed
One-click scans. No signup required.
| Gemma 4 variant | BF16 | SFP8 | Q4_0 |
|---|---|---|---|
| E2B | Approximately 11.4 GB | Approximately 5.7 GB | Approximately 2.9 GB |
| E4B | Approximately 17.9 GB | Approximately 8.9 GB | Approximately 4.5 GB |
| 12B | Approximately 26.7 GB | Approximately 13.4 GB | Approximately 6.7 GB |
Google describes Gemma 4 variants from edge-oriented E2B and E4B models through 12B, 26B A4B, and 31B models aimed at consumer GPUs and workstations. Its model card lists text and image support across the family, with audio for E2B, E4B, and 12B. It lists 128K context for E2B and E4B, and 256K for 12B and 31B; the overview gives 256K for 26B A4B. The 26B A4B is a mixture-of-experts model with 25.2B total parameters and 3.8B active parameters. Active parameters do not mean only that smaller subset must be resident: Google says all 26B parameters must be loaded for fast routing and inference. See Google’s Gemma 4 overview and the Gemma 4 model card for the details of the family and its variants.
Keep local use local—and troubleshoot carefully
Running inference locally can keep the workflow on your computer, but it is not by itself proof that your data stays there. Review the app’s settings and network behavior, including telemetry, extensions, cloud features, and any connected services. Do not expose a local API or server to a public network unless you understand its authentication and access controls.
Quick Recap
- Model will not load: Check free RAM and VRAM, lower the context length, select a smaller or more heavily quantized model, and close other memory-intensive applications.
- Model runs slowly: Check whether the runtime is using the intended accelerator or falling back partly or wholly to CPU. Hybrid inference can work while still being slower than expected.
- Download or file fails: Verify that the runtime supports the model format and that the download completed successfully.
- Unexpected capabilities or terms: Recheck the exact variant, model card, and license rather than relying on a model-family name or catalog label.
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

