Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Ollama lets you download and run Gemma 4 locally, including models that accept image input. The current local lineup has five variants: E2B, E4B, 12B Unified, 26B A4B, and 31B; the untagged gemma4 name currently resolves to E4B.

This guide explains how to install Ollama, pull and run a model, and compare the five variants. Google released the original four on March 31, 2026, then added 12B Unified on June 3, 2026, according to its Gemma release notes.

Current Gemma 4 models in Ollama

These are the current local tags, excluding MLX and cloud variants. Sizes, context windows, and listed input types are from the Ollama Gemma 4 library.

Model Ollama tag Download size Context Input listed by Ollama
Gemma 4 E2B gemma4:e2b 7.2 GB 128K tokens Text, image
Gemma 4 E4B gemma4:e4b 9.6 GB 128K tokens Text, image
Gemma 4 12B Unified gemma4:12b 7.6 GB 256K tokens Text, image
Gemma 4 26B A4B gemma4:26b 18 GB 256K tokens Text, image
Gemma 4 31B gemma4:31b 20 GB 256K tokens Text, image

These are Ollama’s displayed download sizes, not fixed RAM or VRAM requirements. Runtime memory also depends on context length, image inputs, KV-cache usage, hardware backend, and Ollama settings; the official pages do not publish minimum memory figures for each tag.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
ASUS ESC8000A-E13 4U AI GPU Server Barebones with 3+1 3200W Titanimum CRPS Supporting Eight (8) 2-Slot Server GPUs (e.g. Pro 6000, H200), Dual (2) EPYC 9005 CPUs & 24-Channels of DDR5 ECC RDIMM RAM
  • [ Maximum AI Compute Power ] Dominate complex workloads with the ASUS ESC8000A-E13. This 4U rack server is a powerhouse engineered for mass-scale AI, machine learning, and deep training. Featuring support for dual AMD EPYC 9005/9004 processors and up to eight dual-slot GPUs, it delivers the raw computational muscle required to train LLMs and run complex simulations effortlessly. Accelerate your data science pipeline and transform raw data into actionable intelligence faster than ever.
  • [ Advanced Thermal Efficiency ] High performance demands elite cooling. The ESC8000A-E13 features a cutting-edge aerodynamic design with independent CPU and GPU airflow tunnels. Equipped with redundant hot-swap fans and optimized for liquid cooling integrations, this 4U server ensures maximum uptime under heavy, sustained workloads. Keep your data center running cool, quiet, and highly efficient while preventing thermal throttling during mission-critical enterprise operations.
  • [ Scale with Flexible Storage ] Future-proof your infrastructure with unmatched storage and expansion flexibility. This offers comprehensive front-panel drive bays supporting Gen5 NVMe, SAS, or SATA drives alongside multiple PCIe 5.0 slots. Designed as a high-density 4U server capable of housing eight dual-slot GPUs: NVD H200, RTX PRO 6000 Blackwell, RTX PRO 4500 Blackwell or AMD Instinct MI350P PCIe Card, each supporting up to 600 watts.
  • [ Enterprise-Grade Reliability ] Minimize downtime and secure your ecosystem with server-grade redundancy. The ESC8000A-E13 is built for 24/7 continuous operation, boasting 2+2 redundant (3200W total) 80 PLUS Titanium power supplies and integrated ASUS ASMB11-iKVM for comprehensive out-of-band management. Ideal for cloud service providers, rendering farms, and large enterprise infrastructure, it combines robust physical hardware with smart remote monitoring to safeguard your digital assets.
  • [Reliability Guaranteed] Shop with total peace of mind knowing that every new computer component we sell is backed by our EPC 3-year warranty. Whether you are investing in high-speed DDR5 RAM or a powerhouse GPU, we protect your build against defects and performance failures. We stand firmly behind the quality of our hardware, ensuring that your setup remains fast, stable, and secure for years to come.

What the labels mean

  • E2B and E4B: The E means effective parameters. Google’s Gemma 4 model card lists E2B as 2.3B effective parameters, or 5.1B including embeddings, and E4B as 4.5B effective parameters, or 8B including embeddings.
  • 12B Unified: This encoder-free multimodal model projects image patches and audio waveforms directly into the language model’s embedding space, according to Google’s model card.
  • 26B A4B: This mixture-of-experts model has 25.2B total parameters and 3.8B active during inference, according to Google’s model card. Its download size reflects stored weights, not the parameters used for every token.
  • 31B: This is a dense model with approximately 30.7B parameters, according to Google’s model card.

Which Gemma 4 size should you use?

Best choice Model Reason
Smallest local starting point gemma4:e2b Smallest download and a 128K context window; designed for on-device execution.
Compact general-purpose model gemma4:e4b Relatively small and currently the untagged default.
Long context plus native audio gemma4:12b 256K context and native audio support at the Gemma model level.
MoE experimentation gemma4:26b 256K context, with only a subset of parameters active per inference step.
Largest dense option gemma4:31b Largest current dense model, with a 256K context window.

Google’s model card documents native audio support for E2B, E4B, and 12B Unified. It lists 26B A4B and 31B for text and image, not native audio. Ollama’s library currently lists all five local variants as accepting text and image input, so do not assume native audio is available through Ollama’s commands.

Install Ollama

  1. Open Ollama’s Download page.
  2. Select your operating system and click Download.
  3. On Windows, run the downloaded .exe installer.
  4. On macOS, unpack the ZIP and move the Ollama application folder to Applications.
  5. On Linux, follow the Bash installer instructions shown for the platform.

Ollama does not install a model by default. You must pull Gemma 4 separately. These installation steps and the version-check example are documented in Google’s Ollama integration guide.

Open a new terminal and check that the command is available:

ollama –version

A successful result resembles: ollama version is #.#.##

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

If the shell reports that ollama cannot be found, the executable may not be on the operating system’s PATH. Reopen the terminal and check the installer or application setup.

Download Gemma 4

To download the current default, run:

ollama pull gemma4

The untagged name currently resolves to gemma4:latest, listed at 9.6 GB in the Ollama library, matching E4B. Use an explicit tag when you want scripts and notes to identify the model unambiguously.

Rank #2
Sale
HPE NVIDIA Tesla V100 32GB HBM2 PCIe 3.0 x16 Passive GPU Computational Accelerator for AI Machine Learning HPC Deep Learning 699-2G500-0216-400 (Renewed)
  • NVIDIA Volta GV100 Architecture — 4,608 CUDA Cores, 640 1st-Gen Tensor Cores delivering 14 TFLOPS FP32 and 112 TFLOPS deep learning performance for AI training, inference, HPC, and scientific computing workloads
  • 32GB HBM2 ECC Memory — 900 GB/s Bandwidth — High-bandwidth memory on a 4096-bit bus with ECC error correction provides the memory capacity and throughput required for the largest AI models, simulations, and datasets
  • PCIe 3.0 x16 Interface — 250W TDP — Standard PCIe Gen3 connectivity with passive cooling designed for enterprise rack server deployment in HPE ProLiant, Dell PowerEdge, and Supermicro platforms with adequate chassis airflow
  • NVLink — Scale to 96GB Unified Memory — Connect two V100 GPUs via NVLink at 300 GB/s bi-directional bandwidth to scale GPU memory from 32GB to 96GB for larger AI training and HPC workloads
  • Multi-Precision Computing — Supports FP64 (7 TFLOPS), FP32 (14 TFLOPS), FP16 (112 TFLOPS) and INT8 precision modes for flexible deployment across training, inference, and scientific simulation workloads

To download a particular variant, run the command for that tag:

  • ollama pull gemma4:e2b
  • ollama pull gemma4:e4b
  • ollama pull gemma4:12b
  • ollama pull gemma4:26b
  • ollama pull gemma4:31b

Downloading all five requires roughly 62.4 GB of model storage, based on the sizes in Ollama’s library, before allowing for additional runtime data. You do not need all five installed.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Check installed models

List the models stored locally:

ollama list

Ollama model names use the form model_name:tag. If a model is missing from the list, pull it before running it.

Run Gemma 4 from the terminal

Start an interactive session with a downloaded tag. For example:

ollama run gemma4:e2b

Replace e2b with e4b, 12b, 26b, or 31b to run another variant. After the model starts, type a prompt. For a one-shot prompt using the untagged default, run:

ollama run gemma4 “roses are red”

This uses the current E4B default. To choose a model explicitly, run:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #3
Rosewill 4U Server Chassis Case|Supports up to 4 GPUs|8 Hot-Swap 3.5"/2.5" SATA/SAS up to 12Gbps|E-ATX Compatible|3x 12038 Hot-Swap Fans,2 Rear 8038 Fans|USB 3.2 Type-C|With Rail Kit-RSV-AI01
  • AI-Optimized: Designed to support up to 4 GPUs, it is perfect for handling intensive AI and machine learning tasks, ensuring high performance and scalability for advanced computational needs.
  • Intelligent Storage: Equipped with 8 hot-swappable 3.5" SATA/SAS drives (12Gbps), featuring SGPIO and temperature control, it ensures efficient data management and reliable storage performance.
  • Robust Cooling: The system includes 3x 12038 hot-swap PWM fans and 2x 8038 rear fans, providing advanced thermal management to maintain optimal temperatures and ensure stable operation under heavy workloads.
  • Rack-Ready: Comes with a pre-installed rail kit, allowing for quick and easy installation in standard 19-inch server racks, making it ideal for data center environments and enterprise setups.
  • Versatile Connectivity: Offers USB 3.0 and the latest USB 3.2 Type-C ports, ensuring high-speed data transfer and compatibility with a wide range of peripherals and devices for enhanced connectivity options.

ollama run gemma4:12b “Explain how a reverse proxy works”

Use an image

Ollama’s current Gemma 4 variants accept image input. Google’s documented command-line form is:

ollama run gemma4 “caption this image /Users/$USER/Desktop/surprise.png”

Use a path appropriate to your operating system. The image path must point to a file the Ollama process can read.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

At the Gemma model level, E2B, E4B, and 12B Unified also support audio. However, Ollama’s library lists its local Gemma 4 entries as Text, Image, so do not assume every native Gemma capability is available through the same Ollama command syntax.

Call Gemma 4 through Ollama’s local API

When Ollama is running, its local generation endpoint is:

Rank #4
ASRock Radeon AI PRO R9700 Creator 32GB Professional Graphics Card, 2920 MHz Boost Clock, GDDR6, AMD RDNA 4, AI-Accelerators, DisplayPort 2.1a, PCIe 5.0, Blower Cooler
  • Professional AI & Creator Workstation: AMD Radeon AI PRO R9700 GPU with 32GB GDDR6 is engineered for AI development, professional content creation, and compute-intensive workloads.
  • Massive 32GB Memory Capacity: 32GB of GDDR6 memory on a 256-bit bus provides ample bandwidth for large AI models, 8K video editing, and complex 3D rendering.
  • Advanced RDNA 4 with AI Accelerators: 64 Compute Units with 3rd Gen Ray Tracing and dedicated 2nd Gen AI Accelerators for groundbreaking AI performance and visual computing.
  • Professional Blower Cooling: Efficient single blower design exhausts heat directly out of the chassis, ideal for multi-GPU workstation and server configurations.
  • Enterprise-Grade Thermal Solution: Vapor chamber heatsink with industrial Honeywell PTM7950 thermal interface material ensures reliable cooling under sustained professional loads.

http://localhost:11434/api/generate

A basic request looks like this:

curl http://localhost:11434/api/generate -d ‘{“model”:”gemma4″,”prompt”:”roses are red”}’

For a specific model, replace gemma4 with a downloaded tag such as gemma4:26b. Image requests add an images array containing base64-encoded image data:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

curl http://localhost:11434/api/generate -d ‘{“model”:”gemma4″,”prompt”:”caption this image”,”images”:[“…”]}’

The endpoint returns generated output as a streamed response by default. Applications integrating it should account for multiple JSON response lines rather than assuming one complete JSON object for the whole generation.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Quantization and local storage

Google’s Ollama integration uses quantized Gemma models in GGUF format. Quantization stores values with less precision, reducing storage and compute requirements, with a possible reduction in output quality. The GB number shown in the Ollama library is the download size of the quantized model, not the complete memory requirement while generating.

Why some guides still say there are four sizes

Google’s Ollama integration guide still describes the original four-model lineup and omits 12B Unified. Google’s release notes record the later June 3 release, and Ollama’s library lists gemma4:12b. For current commands, use the Ollama library tags rather than copying the older four-size list.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Two common mistakes are:

  • Counting gemma4:latest as a fifth model: It is an alias for the current E4B default, not a separate size.
  • Assuming all models support audio: Google documents native audio for E2B, E4B, and 12B, not 26B A4B or 31B.

FAQ

How many Gemma 4 models are currently available in Ollama?

Five local variants are currently listed: E2B, E4B, 12B Unified, 26B A4B, and 31B. The older four-size list predates the 12B release.

What does ollama pull gemma4 download?

The untagged name currently resolves to gemma4:latest, which the Ollama library lists at 9.6 GB, matching E4B. Use ollama pull gemma4:e4b to specify the variant explicitly.

Can Gemma 4 process images locally with Ollama?

Yes. Ollama’s current Gemma 4 listings accept text and image input. Google’s documented example is ollama run gemma4 “caption this image /path/to/image.png”.

Does the 18-GB 26B model require exactly 18 GB of RAM or VRAM?

No. Eighteen GB is the displayed model download size. Runtime memory depends on context length, image inputs, KV-cache usage, backend, and Ollama settings; the cited official pages do not publish a fixed minimum RAM or VRAM figure for this tag.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Why is 26B A4B called a 26B model if only about 4B parameters are active?

It is a mixture-of-experts model. Google’s model card lists 25.2B total parameters and 3.8B active parameters during inference. A4B describes the approximate active-parameter scale, not the total file size.

The Bottom Line

Install Ollama, verify it with ollama –version, then pull the tag that fits the task. Start with gemma4:e2b for the smallest download, use gemma4:e4b for the compact default, or consider gemma4:12b, gemma4:26b, and gemma4:31b when you want a larger context window or model capacity. The current lineup has five sizes, despite older four-size documentation.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.