To send a prompt to a language model running on your computer, install Ollama, download a model, then call Ollama’s local chat API from Python. This walkthrough uses Ollama’s documented gemma4:e2b example and the official Python client. It is a local-development setup, not guidance for exposing a model server to the internet.
What you need for a local LLM API in Python
- A supported macOS, Windows, or Linux computer with Ollama installed.
- Python and a terminal. The documentation used here does not specify a Python version or a tested cross-platform matrix, so use a Python version supported by the current Ollama package.
- Disk space and memory appropriate for the model you choose.
Ollama’s current quickstart model example, Gemma 4 E2B, is about 7.2 GB to download. Ollama recommends 8 GB of available VRAM, or unified memory on a Mac, for that example. Those figures apply to this model example, not to every local model. Larger context windows need more memory; with less VRAM, Ollama may use system RAM, which can make responses slower. See the Ollama quickstart for current installation options and model guidance.
Install Ollama and download a model
- Install Ollama for your operating system using the official download page. Open the Ollama app where applicable, or follow the terminal setup shown for your platform.
- Open a terminal and download the model from the quickstart example:
ollama pull gemma4:e2b - Wait for the download to finish. Model identifiers may change as the library evolves; check the current quickstart if this name is unavailable.
A model name can include a tag, such as :e2b. The API reference says the tag is optional and defaults to latest; using an explicit tag makes it clear which model identifier your example requests. See Ollama’s API reference.
Start or confirm the local server
Ollama provides a local API at http://localhost:11434/api. On macOS and Windows, the app typically runs the service after launch. The quickstart specifically instructs Linux users to start the server with this command if it is not already running:
Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Fix the driver behind crashes, sound loss and screen glitches3Clear out junk files and repair common Windows errors#1 Best Overall
- Includes Raspberry Pi 5 with 2.4Ghz 64-bit quad-core CPU (8GB RAM)
- Includes 128GB Micro SD Card pre-loaded with 64-bit Raspberry Pi OS, USB MicroSD Card Reader
- CanaKit Turbine Black Case for the Raspberry Pi 5
- CanaKit Low Noise Bearing System Fan
- Mega Heat Sink - Black Anodized
ollama serve
Keep that terminal open while you use the API. If the server is already active, starting another instance is unnecessary. Local requests do not need an API key, as Ollama states in its API introduction.
Send a first local LLM API request
The chat endpoint is http://localhost:11434/api/chat. A chat request supplies the model and a list of role-tagged messages. Set stream to false to receive one response object rather than a stream of objects.
Rank #2
- Includes Raspberry Pi 4 4GB Model B with 1.5GHz 64-bit quad-core CPU (4GB RAM)
- Includes Pre-Loaded 32GB EVO+ Micro SD Card (Class 10), USB MicroSD Card Reader
- CanaKit Premium High-Gloss Raspberry Pi 4 Case with Integrated Fan Mount, CanaKit Low Noise Bearing System Fan
- CanaKit 3.5A USB-C Raspberry Pi 4 Power Supply (US Plug) with Noise Filter, Set of Heat Sinks, Display Cable - 6 foot (Supports up to 4K60p)
- CanaKit USB-C PiSwitch (On/Off Power Switch for Raspberry Pi 4)
curl http://localhost:11434/api/chat -d '{
"model": "gemma4:e2b",
"messages": [
{"role": "user", "content": "Explain what a local API does in one sentence."}
],
"stream": false
}'
The response contains a message object; its content field is the model’s reply. The API reference describes the request and response behavior, including streaming, at Ollama’s API reference.
Run an LLM locally with Python
Ollama’s Python README documents installing the client with pip, calling chat, and reading response.message.content. Create a project folder, install the package, and save the following as chat.py:
The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Rank #3
- Not including the Raspberry Pi 5 (8GB), the Crowpi advanced version comes with the Raspberry Pi 5
- ELECROW Black Case for the Raspberry Pi 5, CrowPi is equipped with a 9-inch HD touchscreen along with a camera; All the regular components used in DIY electronics are packed into the CrowPi development board, such as LCD, LED matrix, buzzer, light sensor, PIR sensor, ultrasonic sensor, IR sensor, etc
- Raspberry Pi Sensors: The Crowpi raspberry pi 5 programming kit is jam-packed with lots of buttons such as 19 different sensors in a tidy easy to use package; You don't have to wait and wire things
- Build Quality: Solid ABS shell and well made components in one place make it strong and convenient to travel
- Programming Lessons: This raspberry pi 5 learning kit ships with step by step instructions and provides 21 lessons to take you through identifying components reading code and running it in the terminal
python -m pip install ollama
from ollama import chat
response = chat(
model="gemma4:e2b",
messages=[
{"role": "user", "content": "Explain what a local API does in one sentence."}
],
)
print(response.message.content)
Run it in the same Python environment where you installed the package:
python chat.py
The call waits for the response, then prints its text. This follows the official Python example; it is not a claim of independent testing. See the Ollama Python repository README for the client’s current instructions.
Rank #4
- Fully assembled for plug-and-play operation
- Includes Raspberry Pi 5 with 8GB RAM
- 256 GB PCIe Pi NVMe SSD (Pre-loaded with Pi 64-Bit OS)
- M.2 HAT+
- CanaKit Turbine Black Case for the Pi 5
Choose between the Ollama and OpenAI-compatible Python clients
If you already use the OpenAI Python client, Ollama also provides an OpenAI-compatible API at http://localhost:11434/v1. The quickstart demonstrates the chat completions endpoint. The two documented paths differ as follows:
| Path | Endpoint | Where response text appears | Coverage |
|---|---|---|---|
| Ollama Python client | Ollama /api/chat endpoint |
response.message.content |
Uses Ollama’s Python client; see the official README. |
| OpenAI-compatible client | http://localhost:11434/v1/chat/completions |
choices[0].message.content |
Compatibility covers a subset of the original OpenAI API. |
The compatibility endpoint does not require a local API key. Ollama’s quickstart shows the request pattern and the choices[0].message.content response path; the API introduction documents the base URL. Read the quickstart and API introduction before adapting code that relies on other OpenAI features.
Best Value
- 【What you Get】You will get 1*Pi 5 8GB Single Board,1*RasTech Case,1*Active Cooler,1*Screwdriver,1*Installation instructions,12-month free warranty, lifetime service, 24-hour prompt and friendly response.
- 【More Connectors】There are two USB 3.0 ports(5Gbps simultaneously) and two USB 2.0 ports, which triple total bandwidth ,support any combination of up to two cameras or displays. Peak SD card performance is doubled through support for the SDR104 high-speed mode. It provides a smooth desktop experience for you. Offer Gigabit Ethernet and a PCIe interface, along with dual-band Wi-Fi and Bluetooth 5.0/BLE wireless capability. The RasTech Pi 5 Kit use the new 27W 5.1V 5A USB-C power connector.
- 【 Support Dual 4Kp60 Display 】Each of the two microHDMI sockets can control a 4K display at 60 Hertz, now support HDR, offering super HD video for media streaming projects. RPi 5 is the first RPi model that comes with a PCI Express port (PCIe 2.0 x1 with 500 MB/s) to attach SSDs (requires separate M.2 HAT).
- 【 Excellent Chips And Applications】Pi 5 is a full-size Pi computer using silicon built in-house at Pi. The RP1 “southbridge” provides the bulk of the I/O capabilities for Pi 5. Pi 5 is more friendly and convenient in the development of Internet of Things, Web development, machine identification, automatic control and other electronic equipment applications and network.
- 【 Faster CPU, Better GPU 】 Pi 5 features a Broadcom BCM2712 64-bit quad-core Arm Cortex-A76 processor running at 2.4GHz, it delivers a 2–3× increase in CPU performance relative to RaspberryPi 4. The 800MHz VideoCore VII GPU is compatible to OpenGL ES 3.1 and Vulkan 1.2, substantial uplift in graphics performance. Pi 5 Offers lightning-fast CPU speed, a PCI Express interface, a Real Time Clock (RTC) and a power button and runs significantly cooler than Pi 4.
What local means—and what it does not guarantee
This project sends requests to a server on your own computer. That does not, by itself, establish that every configuration is private or secure if you expose the service beyond your machine. The documented setup here is for local development; do not treat it as production deployment guidance.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

