Ollama lets you download and run Gemma 4 locally, including models that accept image input. The current local lineup has five variants: E2B, E4B, 12B Unified, 26B A4B, and 31B; the untagged gemma4 name currently resolves to E4B.
This guide explains how to install Ollama, pull and run a model, and compare the five variants. Google released the original four on March 31, 2026, then added 12B Unified on June 3, 2026, according to its Gemma release notes.
Current Gemma 4 models in Ollama
These are the current local tags, excluding MLX and cloud variants. Sizes, context windows, and listed input types are from the Ollama Gemma 4 library.
| Model | Ollama tag | Download size | Context | Input listed by Ollama |
|---|---|---|---|---|
| Gemma 4 E2B | gemma4:e2b | 7.2 GB | 128K tokens | Text, image |
| Gemma 4 E4B | gemma4:e4b | 9.6 GB | 128K tokens | Text, image |
| Gemma 4 12B Unified | gemma4:12b | 7.6 GB | 256K tokens | Text, image |
| Gemma 4 26B A4B | gemma4:26b | 18 GB | 256K tokens | Text, image |
| Gemma 4 31B | gemma4:31b | 20 GB | 256K tokens | Text, image |
These are Ollama’s displayed download sizes, not fixed RAM or VRAM requirements. Runtime memory also depends on context length, image inputs, KV-cache usage, hardware backend, and Ollama settings; the official pages do not publish minimum memory figures for each tag.
#1 Best Overall
- [ Maximum AI Compute Power ] Dominate complex workloads with the ASUS ESC8000A-E13. This 4U rack server is a powerhouse engineered for mass-scale AI, machine learning, and deep training. Featuring support for dual AMD EPYC 9005/9004 processors and up to eight dual-slot GPUs, it delivers the raw computational muscle required to train LLMs and run complex simulations effortlessly. Accelerate your data science pipeline and transform raw data into actionable intelligence faster than ever.
- [ Advanced Thermal Efficiency ] High performance demands elite cooling. The ESC8000A-E13 features a cutting-edge aerodynamic design with independent CPU and GPU airflow tunnels. Equipped with redundant hot-swap fans and optimized for liquid cooling integrations, this 4U server ensures maximum uptime under heavy, sustained workloads. Keep your data center running cool, quiet, and highly efficient while preventing thermal throttling during mission-critical enterprise operations.
- [ Scale with Flexible Storage ] Future-proof your infrastructure with unmatched storage and expansion flexibility. This offers comprehensive front-panel drive bays supporting Gen5 NVMe, SAS, or SATA drives alongside multiple PCIe 5.0 slots. Designed as a high-density 4U server capable of housing eight dual-slot GPUs: NVD H200, RTX PRO 6000 Blackwell, RTX PRO 4500 Blackwell or AMD Instinct MI350P PCIe Card, each supporting up to 600 watts.
- [ Enterprise-Grade Reliability ] Minimize downtime and secure your ecosystem with server-grade redundancy. The ESC8000A-E13 is built for 24/7 continuous operation, boasting 2+2 redundant (3200W total) 80 PLUS Titanium power supplies and integrated ASUS ASMB11-iKVM for comprehensive out-of-band management. Ideal for cloud service providers, rendering farms, and large enterprise infrastructure, it combines robust physical hardware with smart remote monitoring to safeguard your digital assets.
- [Reliability Guaranteed] Shop with total peace of mind knowing that every new computer component we sell is backed by our EPC 3-year warranty. Whether you are investing in high-speed DDR5 RAM or a powerhouse GPU, we protect your build against defects and performance failures. We stand firmly behind the quality of our hardware, ensuring that your setup remains fast, stable, and secure for years to come.
What the labels mean
- E2B and E4B: The E means effective parameters. Google’s Gemma 4 model card lists E2B as 2.3B effective parameters, or 5.1B including embeddings, and E4B as 4.5B effective parameters, or 8B including embeddings.
- 12B Unified: This encoder-free multimodal model projects image patches and audio waveforms directly into the language model’s embedding space, according to Google’s model card.
- 26B A4B: This mixture-of-experts model has 25.2B total parameters and 3.8B active during inference, according to Google’s model card. Its download size reflects stored weights, not the parameters used for every token.
- 31B: This is a dense model with approximately 30.7B parameters, according to Google’s model card.
Which Gemma 4 size should you use?
| Best choice | Model | Reason |
|---|---|---|
| Smallest local starting point | gemma4:e2b | Smallest download and a 128K context window; designed for on-device execution. |
| Compact general-purpose model | gemma4:e4b | Relatively small and currently the untagged default. |
| Long context plus native audio | gemma4:12b | 256K context and native audio support at the Gemma model level. |
| MoE experimentation | gemma4:26b | 256K context, with only a subset of parameters active per inference step. |
| Largest dense option | gemma4:31b | Largest current dense model, with a 256K context window. |
Google’s model card documents native audio support for E2B, E4B, and 12B Unified. It lists 26B A4B and 31B for text and image, not native audio. Ollama’s library currently lists all five local variants as accepting text and image input, so do not assume native audio is available through Ollama’s commands.
Install Ollama
- Open Ollama’s Download page.
- Select your operating system and click Download.
- On Windows, run the downloaded .exe installer.
- On macOS, unpack the ZIP and move the Ollama application folder to Applications.
- On Linux, follow the Bash installer instructions shown for the platform.
Ollama does not install a model by default. You must pull Gemma 4 separately. These installation steps and the version-check example are documented in Google’s Ollama integration guide.
Open a new terminal and check that the command is available:
ollama –version
A successful result resembles: ollama version is #.#.##
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Clear out junk files and repair common Windows errors3Scan for outdated or missing drivers - takes under a minuteIf the shell reports that ollama cannot be found, the executable may not be on the operating system’s PATH. Reopen the terminal and check the installer or application setup.
Download Gemma 4
To download the current default, run:
ollama pull gemma4
The untagged name currently resolves to gemma4:latest, listed at 9.6 GB in the Ollama library, matching E4B. Use an explicit tag when you want scripts and notes to identify the model unambiguously.
Rank #2
- NVIDIA Volta GV100 Architecture — 4,608 CUDA Cores, 640 1st-Gen Tensor Cores delivering 14 TFLOPS FP32 and 112 TFLOPS deep learning performance for AI training, inference, HPC, and scientific computing workloads
- 32GB HBM2 ECC Memory — 900 GB/s Bandwidth — High-bandwidth memory on a 4096-bit bus with ECC error correction provides the memory capacity and throughput required for the largest AI models, simulations, and datasets
- PCIe 3.0 x16 Interface — 250W TDP — Standard PCIe Gen3 connectivity with passive cooling designed for enterprise rack server deployment in HPE ProLiant, Dell PowerEdge, and Supermicro platforms with adequate chassis airflow
- NVLink — Scale to 96GB Unified Memory — Connect two V100 GPUs via NVLink at 300 GB/s bi-directional bandwidth to scale GPU memory from 32GB to 96GB for larger AI training and HPC workloads
- Multi-Precision Computing — Supports FP64 (7 TFLOPS), FP32 (14 TFLOPS), FP16 (112 TFLOPS) and INT8 precision modes for flexible deployment across training, inference, and scientific simulation workloads
To download a particular variant, run the command for that tag:
- ollama pull gemma4:e2b
- ollama pull gemma4:e4b
- ollama pull gemma4:12b
- ollama pull gemma4:26b
- ollama pull gemma4:31b
Downloading all five requires roughly 62.4 GB of model storage, based on the sizes in Ollama’s library, before allowing for additional runtime data. You do not need all five installed.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Check installed models
List the models stored locally:
ollama list
Ollama model names use the form model_name:tag. If a model is missing from the list, pull it before running it.
Run Gemma 4 from the terminal
Start an interactive session with a downloaded tag. For example:
ollama run gemma4:e2b
Replace e2b with e4b, 12b, 26b, or 31b to run another variant. After the model starts, type a prompt. For a one-shot prompt using the untagged default, run:
ollama run gemma4 “roses are red”
This uses the current E4B default. To choose a model explicitly, run:
Quick wins for a faster PC:
Repair Windows errors before they cause bigger problemsFix Now →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Clear out junk files and repair common Windows errorsFree Scan →Rank #3
- AI-Optimized: Designed to support up to 4 GPUs, it is perfect for handling intensive AI and machine learning tasks, ensuring high performance and scalability for advanced computational needs.
- Intelligent Storage: Equipped with 8 hot-swappable 3.5" SATA/SAS drives (12Gbps), featuring SGPIO and temperature control, it ensures efficient data management and reliable storage performance.
- Robust Cooling: The system includes 3x 12038 hot-swap PWM fans and 2x 8038 rear fans, providing advanced thermal management to maintain optimal temperatures and ensure stable operation under heavy workloads.
- Rack-Ready: Comes with a pre-installed rail kit, allowing for quick and easy installation in standard 19-inch server racks, making it ideal for data center environments and enterprise setups.
- Versatile Connectivity: Offers USB 3.0 and the latest USB 3.2 Type-C ports, ensuring high-speed data transfer and compatibility with a wide range of peripherals and devices for enhanced connectivity options.
ollama run gemma4:12b “Explain how a reverse proxy works”
Use an image
Ollama’s current Gemma 4 variants accept image input. Google’s documented command-line form is:
ollama run gemma4 “caption this image /Users/$USER/Desktop/surprise.png”
Use a path appropriate to your operating system. The image path must point to a file the Ollama process can read.
At the Gemma model level, E2B, E4B, and 12B Unified also support audio. However, Ollama’s library lists its local Gemma 4 entries as Text, Image, so do not assume every native Gemma capability is available through the same Ollama command syntax.
Call Gemma 4 through Ollama’s local API
When Ollama is running, its local generation endpoint is:
Rank #4
- Professional AI & Creator Workstation: AMD Radeon AI PRO R9700 GPU with 32GB GDDR6 is engineered for AI development, professional content creation, and compute-intensive workloads.
- Massive 32GB Memory Capacity: 32GB of GDDR6 memory on a 256-bit bus provides ample bandwidth for large AI models, 8K video editing, and complex 3D rendering.
- Advanced RDNA 4 with AI Accelerators: 64 Compute Units with 3rd Gen Ray Tracing and dedicated 2nd Gen AI Accelerators for groundbreaking AI performance and visual computing.
- Professional Blower Cooling: Efficient single blower design exhausts heat directly out of the chassis, ideal for multi-GPU workstation and server configurations.
- Enterprise-Grade Thermal Solution: Vapor chamber heatsink with industrial Honeywell PTM7950 thermal interface material ensures reliable cooling under sustained professional loads.
http://localhost:11434/api/generate
A basic request looks like this:
curl http://localhost:11434/api/generate -d ‘{“model”:”gemma4″,”prompt”:”roses are red”}’
For a specific model, replace gemma4 with a downloaded tag such as gemma4:26b. Image requests add an images array containing base64-encoded image data:
Recommended Free Tools
curl http://localhost:11434/api/generate -d ‘{“model”:”gemma4″,”prompt”:”caption this image”,”images”:[“…”]}’
The endpoint returns generated output as a streamed response by default. Applications integrating it should account for multiple JSON response lines rather than assuming one complete JSON object for the whole generation.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Quantization and local storage
Google’s Ollama integration uses quantized Gemma models in GGUF format. Quantization stores values with less precision, reducing storage and compute requirements, with a possible reduction in output quality. The GB number shown in the Ollama library is the download size of the quantized model, not the complete memory requirement while generating.
Why some guides still say there are four sizes
Google’s Ollama integration guide still describes the original four-model lineup and omits 12B Unified. Google’s release notes record the later June 3 release, and Ollama’s library lists gemma4:12b. For current commands, use the Ollama library tags rather than copying the older four-size list.
The Tool Desk
Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Best Value
Two common mistakes are:
- Counting gemma4:latest as a fifth model: It is an alias for the current E4B default, not a separate size.
- Assuming all models support audio: Google documents native audio for E2B, E4B, and 12B, not 26B A4B or 31B.
FAQ
How many Gemma 4 models are currently available in Ollama?
Five local variants are currently listed: E2B, E4B, 12B Unified, 26B A4B, and 31B. The older four-size list predates the 12B release.
What does ollama pull gemma4 download?
The untagged name currently resolves to gemma4:latest, which the Ollama library lists at 9.6 GB, matching E4B. Use ollama pull gemma4:e4b to specify the variant explicitly.
Can Gemma 4 process images locally with Ollama?
Yes. Ollama’s current Gemma 4 listings accept text and image input. Google’s documented example is ollama run gemma4 “caption this image /path/to/image.png”.
Does the 18-GB 26B model require exactly 18 GB of RAM or VRAM?
No. Eighteen GB is the displayed model download size. Runtime memory depends on context length, image inputs, KV-cache usage, backend, and Ollama settings; the cited official pages do not publish a fixed minimum RAM or VRAM figure for this tag.
Why is 26B A4B called a 26B model if only about 4B parameters are active?
It is a mixture-of-experts model. Google’s model card lists 25.2B total parameters and 3.8B active parameters during inference. A4B describes the approximate active-parameter scale, not the total file size.
The Bottom Line
Install Ollama, verify it with ollama –version, then pull the tag that fits the task. Start with gemma4:e2b for the smallest download, use gemma4:e4b for the compact default, or consider gemma4:12b, gemma4:26b, and gemma4:31b when you want a larger context window or model capacity. The current lineup has five sizes, despite older four-size documentation.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

