Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

There is no single best self-hosted LLM for every job. Choose a model that fits your task, hardware, context needs, and expected number of users, then select an inference runtime suited to how you plan to run it. This guide explains the choices and gives you a reproducible deployment and evaluation process; it does not claim personal deployment tests or results.

What “self-hosted LLM” means

A self-hosted LLM runs on hardware you control—such as a workstation or server—rather than sending inference requests to a hosted model service. The model’s weights and the software used to run them are separate choices: an open-weight model is not automatically open-source, and its license may impose conditions on use or distribution. Review the specific model’s license and terms before deploying it, especially for commercial or public-facing applications.

There are two decisions: which model fits your workload, and which runtime can serve it reliably on your hardware. A model library listing establishes that an artifact is available; it does not establish its quality, speed, or suitability for your use case.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Which self-hosted LLM should you choose?

Start with the work you need done rather than a universal leaderboard. Ollama’s current library includes model families such as Gemma 4 and Qwen 3.5, but catalog availability and variants can change. Validate a specific model revision against your own representative prompts and constraints.

#1 Best Overall
GMKtec EVO-X2 AI Mini PC Ryzen Al Max+ 395 Superchip 128GB LPDDR5X 2TB SSD
  • EVOLUTION RYZEN AI MAX+ 395 MINI PC - GMKtec EVO-X2 is the next evolution in AI mini PC Ryzen Strix Halo series. Thanks to AMD Simultaneous Multithreading (SMT) the core-count is effectively doubled, to 32 threads. Ryzen AI Max+ 395 has 64 MB of L3 cache and can boost up to 5.1 GHz, depending on the workload. The Ryzen AI Max+ 395 is currently rated as the "most powerful x86 APU" on the market for AI computing.
  • AI NPU with XDNA 2 ARCHITECTURE - Powered by 16 “Zen 5” CPU cores, 50+ peak AI TOPS XDNA 2 NPU and a truly massive integrated GPU driven by 40 AMD RDNA 3.5 CUs, the Ryzen AI MAX+ 395 is a transformative upgrade and delivers a significant performance boost over the competition. The Ryzen AI Max+ 395 excels in consumer AI workloads like the llama.cpp-powered application: LM Studio. Shaping up to be the must-have app for client LLM workloads, LM Studio allows users to locally run the latest language model without any technical knowledge required and unleash their creativity and productivity.
  • AMD RADEON 8090S iGPU GAMING PC - The AMD Radeon RX 8060S offers all 40 CUs with up to 2.9 GHz graphics clock and uses the new RDNA 3.5 architecture. The powerful iGPU is positioned between an RTX 4060 and 4070 laptop GPU and therefore enables gaming in FHD at maximum details in most demanding games. The 8060S can also utilize the full 128GB pool, which is perfect for running LLMs such as Deepseek 70B Q8, which runs comfortably on this machine.
  • EIGHT CHANNEL LPDDR5X - LPDDR5X is a new ground breaking memory small form factor installed on-board. With blazing speeds up to to 8000MT/s, it runs 1.5x faster than the DDR5 SODIMMs; 90% better performance over DDR5 SODIMMs in video conferencing and photo editing; 30% better performance in productivity apps; 12% better performance in digital content workloads.
  • QUAD SCREEN 8K DISPLAY SUPPORT - EVO-X2 AI Mini PC support 4-screen 4K/8K output via HDMI 2.1 (8K@60Hz), DisplayPort 1.4 (4K@60Hz), and dual USB 4 40Gbps Transfer speed (supporting PD3.0/DP1.4/DATA). Ideal for gaming, video editing, and multitasking, it provides expansive and crisp multi-display support.
Workload What to prioritize How to decide
General chat and everyday assistance Useful answers, instruction following, and acceptable latency at a model size your machine can sustain Test common questions, refusal behavior, and longer conversations. A larger variant is not automatically better for your hardware or needs.
Coding Correctness on your languages and repository tasks, ability to follow constraints, and tool-use fit if your workflow needs it Use a fixed set of real coding tasks; check whether suggestions build or tests pass rather than judging only fluency.
Reasoning and multi-step work Reliability across several steps, error handling, and performance at the context length you actually use Include tasks with checkable outcomes and compare the final result, not just the model’s explanation.
Multimodal input or document workflows Support for the required input modality, document length, and any extraction or tool workflow you depend on Confirm the exact model and runtime support the inputs you need; do not infer capability from a family name alone.

Ollama’s Gemma 4 listing shows variants from e2b and e4b to 31B. Its default listing gives a 6.6–9.5 GB artifact-size range, and the larger variants list context windows up to 256K. These are catalog details, not a quality ranking or a complete RAM/VRAM requirement. See the Ollama model library and Gemma 4 listing for the current entries.

Which runtime should you use?

Runtime Best fit Key trade-off
llama.cpp Flexible local inference across CPUs and multiple GPU or vendor backends, including hybrid CPU+GPU operation Its many backend, build, and quantization choices provide flexibility but may require more hands-on configuration.
vLLM Serving workloads where throughput, scheduling, and concurrent requests matter Platform support is release-specific; verify the exact installation path for your hardware and version.
Ollama A straightforward local workflow for discovering and running models from its library Convenience does not remove the need to check model fit, memory use, licensing, or performance on your workload.

Use llama.cpp for backend flexibility

The llama.cpp project documents a C/C++ inference implementation, quantization options from 1.5-bit through 8-bit, and CPU+GPU hybrid inference for models larger than available VRAM. It also documents an OpenAI-compatible server route. This makes it a useful option when you want to tune model representation or mix CPU and GPU resources; check the project documentation for the backend and build instructions applicable to your machine.

Use vLLM when serving throughput is central

vLLM is designed around serving features such as PagedAttention, scheduling, and continuous batching. Its project describes it as “The High-Throughput and Memory-Efficient inference and serving engine for LLMs.” It offers an OpenAI-compatible API. The vLLM v0.31.0 installation documentation, dated May 11, 2026, lists paths for CUDA, ROCm, Intel XPU, and CPU; Apple Silicon support is through the separate vLLM-Metal project. Treat platform compatibility as version-specific and verify the release’s requirements before installing.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use Ollama for a convenient local model workflow

Ollama combines a model library with a local running workflow. Its library can help you find available families and variants, but a listed file size describes the artifact—not the full memory needed during inference. Account for runtime overhead, context/KV cache, and concurrent requests when deciding whether a model will fit.

How to estimate hardware needs

Do not choose hardware from a model’s download size alone. Begin with the exact model artifact and quantization you intend to use, then test the full workload. Actual memory use depends on the model, runtime, context length, and concurrency; the available catalog figures do not establish a universal VRAM chart.

  1. Choose the model and artifact. Record the precise model variant and quantization, not just its family name. Quantization can reduce weight storage, but it is one part of total inference memory.
  2. Set the required context length. Longer prompts and conversations require context/KV cache capacity in addition to model weights. A catalog context-window maximum is not a promise that your hardware can run that maximum comfortably.
  3. Include runtime overhead and concurrency. Leave capacity for the runtime and operating system, and consider how many requests may be active at once. A configuration that works for one interactive user may not work for several simultaneous users.
  4. Test the target backend. Run the exact model and runtime on the intended CPU/GPU combination. With llama.cpp, hybrid CPU+GPU inference is an option for a model exceeding available VRAM, but the resulting speed must be measured on your setup.
  5. Record memory and latency under realistic prompts. Test both ordinary requests and the longest context you expect, then check for slowdowns, out-of-memory errors, or instability before making the deployment permanent.

Ollama reported up to 20% faster NVIDIA performance with Ollama 0.30 in its June 5, 2026 release post, which includes a Gemma 4 26B Q4_K_M test on an NVIDIA RTX 5090. That is a vendor-reported result tied to that release and example, not an independent comparison, a guarantee for other models or systems, or evidence that an RTX 5090 is required. The details are in Ollama’s “Improved performance and model support with GGUF”.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

How to deploy and evaluate a local LLM

A deployment is only meaningful when its configuration and evaluation are recorded. Since hardware, model, runtime, and workload details determine the result, use this checklist to make your own setup reproducible rather than assuming a configuration tested elsewhere will transfer.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  1. Define the workload. Write down the tasks, input types, typical and maximum context, expected response time, and number of simultaneous users.
  2. Select a model variant. Choose a model and quantization that fit the task and likely memory budget. Check that the model’s license and terms permit your intended use.
  3. Choose the runtime. Prefer llama.cpp when backend and quantization flexibility are priorities, vLLM when serving throughput features are central, or Ollama for a convenient library-and-run workflow. Confirm hardware support in the documentation for the version you will install.
  4. Install and record versions. Follow the runtime’s official installation instructions and record the runtime version, operating system, driver or backend details, and model revision or artifact.
  5. Run a fixed evaluation set. Use representative prompts and, where possible, tasks with objectively checkable answers. Keep prompts, settings, context length, and concurrency consistent when comparing configurations.
  6. Measure the result. Record latency, memory use, output quality, errors, and stability under the expected load. Do not describe a model as faster or better without specifying the hardware and test conditions.
  7. Review operational needs. For an API or shared service, check authentication, access controls, logging, updates, and resource limits before exposing it to other users or a network.

A useful deployment log records the operating system; CPU, GPU, and system memory; model name, revision, and quantization; runtime and version; context settings; concurrency; evaluation prompts and method; and observed failures or trade-offs. Without those details, a claim about what was deployed or how it performed is not reproducible.

A practical decision path

  • One person, one local machine: begin with a model and quantization that fit your hardware, then use Ollama for a simple model-library workflow or llama.cpp when you need more backend control.
  • Model is larger than GPU memory: consider a smaller or more heavily quantized artifact, or test llama.cpp’s CPU+GPU hybrid path; compare the resulting latency against your needs.
  • Several users or an API service: evaluate vLLM’s serving features and verify support for the exact platform and release. Measure concurrent workload behavior rather than extrapolating from single-user use.
  • Quality is the deciding factor: compare candidate models on the same task set, hardware, settings, and context requirements. No common independent benchmark in the cited project and catalog material establishes a universal model or runtime winner.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.