iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
To benchmark an LLM on a Raspberry Pi 5, build llama.cpp, run llama-bench against a specific GGUF model, and report prompt-processing and token-generation rates separately. For a repeatable CPU baseline, disable GPU-layer offload with -ngl 0, keep the workload and settings fixed, and save the repeated-run output. Results are specific to the board, model, build, and test conditions—not a universal speed rating for every Pi 5.
What to record before you run a benchmark
A throughput number is useful only when another person can identify what produced it. Record the board and test configuration alongside every result:
- Raspberry Pi 5 memory configuration, operating system, and any relevant power or thermal conditions.
- The exact GGUF filename, model source or revision, and quantization.
- The checked-out
llama.cpprevision, build options, and backend. - Thread count and any changed context, batch, or other benchmark settings.
- Prompt and generation token counts, repetitions, and the output file.
Choose a GGUF artifact supported by your build and confirm that it and its context fit the board’s available memory. No single model size is established as suitable for every Pi 5 memory configuration. For current prerequisites and CMake instructions, use the llama.cpp build guide; project options and defaults can change, so record the revision and options rather than treating an old build command as timeless.
The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Build llama.cpp for the Pi 5
Follow the upstream build guide for the prerequisites and build procedure appropriate to your operating system and desired backend. Keep the build configuration with your benchmark notes: CPU-only and Vulkan-enabled builds are not equivalent test conditions, and a build option that works in one revision may change in another.
#1 Best Overall
- Includes Raspberry Pi 5 with 2.4Ghz 64-bit quad-core CPU (8GB RAM)
- Includes 128GB Micro SD Card pre-loaded with 64-bit Raspberry Pi OS, USB MicroSD Card Reader
- CanaKit Turbine Black Case for the Raspberry Pi 5
- CanaKit Low Noise Bearing System Fan
- Mega Heat Sink - Black Anodized
Run a CPU-only baseline with llama-bench
After building llama.cpp and placing a compatible model file at the indicated path, this example runs prompt processing, generation, a combined test, and five repetitions:
./build/bin/llama-bench
-m models/model.gguf
-ngl 0
-p 512
-n 128
-pg 512,128
-t 4
-r 5
-o jsonl
This is an example using documented benchmark options, not a measurement or a claim that these workload sizes suit every question. Replace the model path with the actual GGUF file. In this command, -ngl 0 requests no GPU-layer offload, -p 512 sets the prompt length, -n 128 sets the generation length, -pg 512,128 requests the combined prompt-plus-generation test, -t 4 sets four threads, -r 5 repeats the tests five times, and -o jsonl selects JSON Lines output. See the llama-bench documentation for the benchmark modes and options.
Rank #2
- Includes Raspberry Pi 5 16GB with 2.4Ghz 64-bit quad-core CPU (16GB RAM)
- Includes 128GB Micro SD Card pre-loaded with 64-bit Raspberry Pi OS, USB MicroSD Card Reader
- CanaKit Turbine Black Case for the Raspberry Pi 5
- CanaKit Low Noise Bearing System Fan
- Mega Heat Sink - Black Anodized
Test prompt processing or generation on its own
Use a prompt-processing test when studying how quickly the model handles input tokens, and a generation test when studying output-token throughput. Set the relevant prompt length with -p or generation length with -n; label the run so its result cannot be mistaken for the other phase. Keep the remaining settings fixed when comparing configurations.
Include context depth when it matters
If testing performance after the KV cache has been populated to a particular depth, record that depth. The benchmark tool documents -d for prefilling the KV cache to a specified depth. Do not compare results with different context depths as if they were the same workload.
Rank #3
- CanaKit Raspberry Pi 5 Essentials Starter Kit
Keep comparisons fair and preserve the output
For a controlled comparison, change one variable at a time. Hold the board, model file, quantization, llama.cpp build, backend, thread count, prompt and generation lengths, context depth, and other benchmark options constant unless one of those is the variable under test. Repeat each run and retain the JSONL output; llama-bench reports average tokens per second and standard deviation.
When reporting a result, separate prompt-processing (pp), text-generation (tg), and combined prompt-plus-generation (pg) measurements where applicable. These represent different parts of inference and should not be collapsed into one unexplained rate. The benchmark documentation also notes that its measurements exclude tokenization and sampling time, so its throughput is not complete application latency.
Rank #4
- All-in-One Complete Kit: This SANOOV RPi 5 bundle comes with Raspberry Pi 5 4GB RAM single board, active cooler, durable ABS case and screwdriver. No extra parts needed, ready to use right out of the box for beginners and hobbyists
- Powerful Single Board Computer: Equipped with 4GB RAM and high-performance processor, delivers fast running speed for 4K playback, AI projects, programming and daily computing tasks. SANOOV for raspberry pi 5 4GB is equipped with broadcom 64 quad-core Arm Cortex A76 processor with gigabit ethernet and upgraded with IEEE 802.11ac Wi-Fi, Bluetooth 5.0 dual-band 2.4Ghz and 5Ghz and Power Over Ethernet (POE). Upgrading delivers 2-3 x speed vs Pi 4, redefining the experience
- Efficient Active Cooler: Effectively lowers operating temperature and prevents performance throttling. Runs quietly even under long-time heavy load, ensures stable operation all day long. SANOOV RPi 5 4GB kit offer an active cooler, which combines an aluminium heatsink with a high-performance PWM fan. Active cooler is fully compatible with the Pi OS, which can effectively reduce the temperature of RPi5 and ensure its good performance during long-term high load operation
- Sturdy ABS Protective Case: Well-fitted for Raspberry Pi 5 board, can be secured with 4 screws to effectively protect the Pi 5 motherboard from damage, reserves full access to all ports and buttons. SANOOV uses ABS material to produce the case, which has a softer texture and feel. Meanwhile, SANOOV case adopts a layered design for easy disassembly and installation. (Tip: The Case cannot install M.2 HAT Add on Board and Solid State Drive!)
- Wide Application & Full Compatibility: Seamlessly compatible with official OS and mainstream peripheral accessories for Raspberry Pi 5. Whether you are a beginner, student, electronics hobbyist or professional developer, this all-in-one kit meets your diverse needs. It excels in IoT projects, robotics design, retro gaming devices, home media servers and other DIY creations. Backed by a large global community, you can easily find guides, technical support and shared projects online
How published Raspberry Pi 5 numbers differ
Published numbers illustrate why workload details belong next to every rate. Raspberry Pi’s September 2026 article reports a llama.cpp Q4_0 result of 24 tokens per second for a setup specifying 1,024 prefill tokens, 256 decode tokens, and four CPU threads. A separate 2026 Pi 5 CPU comparison reports 3.91 tokens per second for a tg64 test and 27.77 tokens per second for pp17; its combined pp17+tg64 run took about 16,998 ms. That report used a Qwen3.5-2B GGUF setup, four threads, and a named llama.cpp build. These measurements use different workloads and model conditions, so neither predicts the speed of another model or configuration. Consult the Raspberry Pi article and the mudler / vllm.cpp benchmark report for their stated setups.
Treat Vulkan acceleration as a separate experiment
A CPU-only run with -ngl 0 provides a clear baseline. Do not assume Raspberry Pi 5 VideoCore/Vulkan offload is available or valid for every build and driver. A 2026 llama.cpp issue describes constraints involving workgroup size and shared memory for the Pi 5 V3D Vulkan path; an earlier issue also records Vulkan problems. Issue reports are cautions, not a complete compatibility matrix.
Best Value
- 【What you Get】You will get 1*Pi 5 8GB Single Board,1*RasTech Case,1*Active Cooler,1*Screwdriver,1*Installation instructions,12-month free warranty, lifetime service, 24-hour prompt and friendly response.
- 【More Connectors】There are two USB 3.0 ports(5Gbps simultaneously) and two USB 2.0 ports, which triple total bandwidth ,support any combination of up to two cameras or displays. Peak SD card performance is doubled through support for the SDR104 high-speed mode. It provides a smooth desktop experience for you. Offer Gigabit Ethernet and a PCIe interface, along with dual-band Wi-Fi and Bluetooth 5.0/BLE wireless capability. The RasTech Pi 5 Kit use the new 27W 5.1V 5A USB-C power connector.
- 【 Support Dual 4Kp60 Display 】Each of the two microHDMI sockets can control a 4K display at 60 Hertz, now support HDR, offering super HD video for media streaming projects. RPi 5 is the first RPi model that comes with a PCI Express port (PCIe 2.0 x1 with 500 MB/s) to attach SSDs (requires separate M.2 HAT).
- 【 Excellent Chips And Applications】Pi 5 is a full-size Pi computer using silicon built in-house at Pi. The RP1 “southbridge” provides the bulk of the I/O capabilities for Pi 5. Pi 5 is more friendly and convenient in the development of Internet of Things, Web development, machine identification, automatic control and other electronic equipment applications and network.
- 【 Faster CPU, Better GPU 】 Pi 5 features a Broadcom BCM2712 64-bit quad-core Arm Cortex-A76 processor running at 2.4GHz, it delivers a 2–3× increase in CPU performance relative to RaspberryPi 4. The 800MHz VideoCore VII GPU is compatible to OpenGL ES 3.1 and Vulkan 1.2, substantial uplift in graphics performance. Pi 5 Offers lightning-fast CPU speed, a PCI Express interface, a Real Time Clock (RTC) and a power button and runs significantly cooler than Pi 4.
If you test Vulkan or another accelerated backend, report the exact llama.cpp revision, Mesa or driver version, build configuration, model, and whether you checked the output for correctness. Keep those results separate from the CPU baseline.
Use a reporting checklist
- Identify the Pi 5 memory configuration, operating system, and thermal and power conditions.
- Name the model source or revision, exact GGUF filename, and quantization.
- Give the llama.cpp revision, build options, backend, and thread count.
- State prompt tokens, generation tokens, context depth, and any changed batch-related settings.
- Report repetitions, average and standard deviation or individual runs, and preserve the raw output.
- Label prompt-processing, generation, and combined rates separately; disclose that tokenization and sampling are excluded.
Only compare a published result after aligning or disclosing these conditions. A matching “tokens per second” label alone does not make two tests comparable.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.
Recommended Free Tools

