To use a quantized model in an application with Ollama, prepare or obtain a compatible quantized model file, import it with a Modelfile, then call the local Ollama API from your app. Importing a GGUF file does not quantize it: Ollama’s import documentation says the file must already be quantized if you want a quantized model.
What quantization means for an Ollama application
Quantization is a way of representing a model’s weights in a selected format and precision. In practice, you choose a model variant that fits your storage and runtime constraints while still producing acceptable results for your application. There is no universally best quantization level: the right choice depends on the model, hardware, prompts, context length, and quality requirements.
GGUF is a model format used by llama.cpp and supported by Ollama’s GGUF import workflow. llama.cpp documents converting model data into GGUF and preparing quantized models with its tools: llama.cpp model documentation.
Prepare the model before importing it
If you start with a GGUF file that is already quantized, you can import that file directly. If you have an unquantized model and want a quantized GGUF, quantize it first with an appropriate tool, such as llama.cpp’s llama-quantize. Ollama’s import documentation explicitly says it does not quantize GGUF models during import: Ollama’s model import guide.
Free tools Windows power users keep installed
One-click scans. No signup required.
#1 Best Overall
Confirm the model’s source, license, architecture and compatibility before using it. A file extension alone does not establish that a model is compatible with your intended Ollama version or application.
Import a GGUF model into Ollama
- Place the model file. Save the compatible GGUF file somewhere accessible to the machine running Ollama.
- Create a Modelfile. In a text file named
Modelfile, pointFROMto the GGUF file. Ollama accepts an absolute path or a path relative to the Modelfile; the Modelfile reference also documents GGUF usage: Modelfile reference.FROM ./ollama-model.gguf - Build the Ollama model. From the directory containing the Modelfile, run:
ollama create my-model -f Modelfile - Smoke-test it locally. Run a brief prompt and confirm that Ollama loads the model and returns a sensible response:
ollama run my-model "Reply with one sentence explaining what this model can do."
For a split GGUF model, Ollama’s import documentation supports using a wildcard path that matches the model’s shards. Keep all required shard files together and follow the filename pattern for that model; do not point FROM at only one shard if the model is split.
Call the local model from an application
Ollama documents http://localhost:11434/api as its local API base and http://localhost:11434/v1 as its OpenAI-compatible local API base. Use the interface that best fits your application and consult the current API documentation for request fields and language-specific examples, since options can evolve: API introduction and Ollama API reference.
Choose a chat endpoint when your application has a sequence of user and assistant messages; choose a generation endpoint when it sends a single prompt and expects a completion. Use the model name you created, such as my-model, as the requested model identifier. For applications that need streaming, structured output, or tool calls, the API documents those capabilities, but the exact behavior depends on the endpoint and whether the model supports the requested capability.
Rank #3
For a local service on the same machine, requests can target the documented localhost base URL. If your application runs on another host or in a container, account for network reachability and service configuration rather than assuming that its own localhost resolves to the Ollama host.
Configure runtime behavior and capacity
A Modelfile can set model behavior and runtime parameters. The reference includes parameters such as num_ctx for context length, temperature for generation variability, and num_predict for output length. Defaults and available settings are version-sensitive, so use the current Modelfile reference rather than relying on a fixed settings recipe.
Rank #4
Memory is a runtime constraint, not just a model-file-size concern. Ollama’s FAQ says available system memory affects concurrent processing, and that context size and parallel request count affect memory needs: Ollama FAQ. Test the context length and concurrency your application will actually use; a model that loads for a short, single request may not fit the same way under longer contexts or simultaneous requests.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Evaluate variants on the workload that matters
Compare quantized variants of the same base model using repeatable prompts and the hardware and settings you expect to deploy. Record task quality, peak system memory or VRAM, latency, throughput, and storage footprint. No universal quality ranking or best quantization level is established for every model and application, so make the decision against your own output requirements.
Best Value
- Quality: Score outputs against a stable set of representative tasks, including important edge cases.
- Latency and throughput: Measure under the same hardware, prompt, context, and concurrency settings.
- Memory: Observe peak use at the intended context size and parallel request count.
- Fit and operations: Confirm the file fits available storage and that the deployed Ollama version supports the model.
Ollama’s June 5, 2026 blog reports “up to 20% faster” NVIDIA performance for Gemma 4 26B on an RTX 5090 using Q4_K_M. That is a vendor-reported result for the named setup, not a general speed guarantee for other models, GPUs, or workloads: Ollama’s GGUF performance and model-support announcement.
Quick Recap
Deployment checklist
- Verify model provenance, license, architecture, and file integrity.
- Confirm the model and GGUF variant work with the Ollama version you plan to deploy.
- Quantize before import if needed; do not expect
ollama createto quantize the file. - Set context and generation parameters deliberately, checking the current Modelfile reference.
- Measure quality, memory, latency, and throughput with representative application inputs.
- Exercise expected concurrency and long-context cases, not only a single short prompt.
- Plan how the application will handle unavailable service, slow responses, and model-loading delays.
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

