Recommended Free Tools
iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
Yes, Ollama can run a small language model on the original Jetson Nano, but the benchmark evidence points to CPU-mode performance rather than dependable GPU acceleration. In K. Kreier’s 2025 test of a 2019 Nano, Ollama 0.6.4 generated TinyLlama 1.1B output at 5.37 tokens per second. A CUDA-enabled llama.cpp build reached 6.28 tokens per second with the same model in that test. Those figures describe one board and setup—not a guarantee for every Nano, model, or software version.
What this benchmark tested
K. Kreier tested the same original Jetson Nano machine from 2019, without overclocking. The comparison included Ollama 0.6.4, released in April 2025, running in CPU mode, as well as llama.cpp builds using CPU execution and CUDA-enabled GPU layers. The benchmark used the prompt “Explain quantum entanglement.”
| # | Preview | Product | Price | |
|---|---|---|---|---|
| 1 |
|
ReComputer J3010-Edge AI Device, NVIDIA Jetson Orin Nano 4GB, 4xUSB 3.2, WiFi/BT, M.2 Key E | $599.00 | Buy on Amazon |
The models included TinyLlama-1.1B-Chat Q4_K_M (669 MB), Gemma 3 1B Q4_K_M (806 MB), and Gemma 3 4B variants. The author’s stated focus is direct: “The main metric to compare here is the token generation.” The test is a useful snapshot of those configurations, not a controlled comparison across every Nano revision or software release.
How fast was Ollama on the Nano?
For TinyLlama 1.1B, the benchmark table lists 5.37 generated tokens per second for Ollama 0.6.4 in CPU mode. Its CUDA-enabled llama.cpp comparison lists 6.28 tokens per second. Kreier describes the latter as around 20% faster for token generation in that comparison. These are the author’s results for the tested model, prompt, board, and software builds; they should not be read as expected speeds for other workloads.
#1 Best Overall
- Brilliant AI Performance for production: The reComputer J3010 is equipped with the same NVIDIA Jetson Orin Nano 5GB production module. You can perform a self - upgrade to Jetpack 6.2. Once upgraded, you'll instantly experience a significant boost in computing power, with the performance leaping from 20 Tops to 34 Tops, offering capabilities comparable to those of the NVIDIA Jetson Orin Nano Super Developer Kit.
- Hand-size edge AI device: compact size at 130mm x120mm x 58.5mm, includes NVIDIA Jetson Orin Nano 4GB production module, a heatsink, enclosure, and a power adapter. Support desktop, wall mount, fit in anywhere
- Expandable with rich I/Os: 4x USB3.2, HDMI 2.1, 2xCSI, 1xRJ45 for GbE, M.2 Key E, M.2 Key M, CAN and GPIO
- Accelerate solution to market: pre-installed Jetpack with NVIDIA JetPack 5.1.1 on the included 128GB NVMe SSD, Linux OS BSP, 128GB SSD, WiFi BT combo module, Antennas x2, support Jetson software and leading AI frameworks and software platforms
- Comprehensive certificates: FCC, CE, RoHS, UKCA
The comparison suggests that CUDA-enabled llama.cpp can improve generation speed for this particular test, but Ollama’s reported result does not demonstrate GPU acceleration. The benchmark’s separate description of GPU load and power applies to GPU-enabled llama.cpp behavior, not to the Ollama CPU-mode run.
Model size, memory, and GPU offload limits
Model file size alone does not tell you whether a model will load or perform well. Runtime memory use and the chosen quantization also matter on a board with constrained resources.
| Test case | Reported result | How to interpret it |
|---|---|---|
| TinyLlama 1.1B Q4_K_M | 5.37 tokens/s with Ollama CPU mode; 6.28 tokens/s with CUDA-enabled llama.cpp | Generation results from the benchmark’s specific TinyLlama test. |
| Gemma 3 1B Q4_K_M | 1.9 GB reported Jetson memory use | Memory figure reported by Kreier for the tested configuration. |
| Gemma 3 4B Q4 variant | 2.8 GB reported Jetson memory use | Memory figure reported by Kreier for the tested configuration; larger variants encountered practical limits or failures. |
Kreier reports that full GPU offload succeeded only for the Q2 quantization attempt in one llama.cpp 4B test. In the described 4B case, Ollama 0.6.4 ran in CPU mode through some tested variants and crashed with Google’s version. This is evidence of limits in those particular combinations, not proof that every 4B model or quantization will behave identically.
Does the original Jetson Nano have current Ollama support?
Support guidance for newer Jetson devices does not establish support for the original Nano. NVIDIA’s Jetson AI Lab Ollama tutorial documents native and Docker installation options for its listed supported hardware, which includes Orin-family devices rather than the original Jetson Nano. Its general performance context is not a benchmark of Ollama on the original Nano.
An Ollama GitHub issue titled “Add binary support for Nvidia Jetson Nano- JetPack 4” records user reports of CPU-only behavior on JetPack 4 and describes CUDA 10 and compiler-toolchain complications. Those reports are historical issue discussion, not a definitive statement of current Ollama policy or compatibility for every installation.
Likewise, NVIDIA’s JetPack SDK setup guidance for the Jetson Orin Nano Developer Kit concerns Orin Nano. JetPack includes accelerated libraries and developer tools, but that newer platform’s setup instructions should not be treated as proof of software compatibility on the original Nano.
Quick Recap
What the results mean if you plan to try it
- Keep the hardware generation explicit: the benchmark is for the original 2019 Jetson Nano, not the Orin Nano.
- Start with a small, quantized model such as the tested TinyLlama 1.1B configuration rather than assuming a larger model will fit or load.
- Check whether the runtime is actually using CPU execution or GPU layers; the 5.37-token/s Ollama figure is a CPU-mode result.
- Compare speed only when the model, prompt, software version, and execution mode are alike. The cited values are not directly comparable with results from newer devices or different models.
- Expect compatibility to depend on the Nano’s JetPack and CUDA toolchain as well as the Ollama or llama.cpp build. The historical reports do not settle every current configuration.
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

