What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
In one Lenovo Yoga 9 15IMH5, a GeForce GTX 1650 Ti Max-Q decoded Gemma 4 E2B at a median 4.14 times the rate of the laptop’s six-core Intel Core i7-10750H. The same test reported a 3.42× GPU lead in prompt prefill and a 3.62× lead end to end. Those are results for this laptop, model, software build, and test protocol—not a general rule that laptop GPUs are always four times faster than CPUs.
What did the ABBA re-test find?
The GPU led on all three reported measures: generating output tokens (decode), processing the prompt before generation (prefill), and the combined end-to-end interval. Decode is especially relevant when judging how quickly a model produces an answer; prefill matters when a user submits a prompt, and end-to-end performance reflects both stages together.
| Measure | Reported GPU advantage | What it describes |
|---|---|---|
| Decode | 4.14× median | Output-token generation rate |
| Prefill | 3.42× median | Processing the input prompt |
| End to end | 3.62× median | Combined prompt processing and generation |
The eight GPU-to-CPU decode ratios in the report ranged from 3.93× to 4.30×. Across the result grid, the CPU decoded at 16.10–18.31 tokens per second and the GPU at 68.22–71.89 tokens per second. These are measurements reported by xbill in a DEV Community article published in 2026, not an independently measured industry statistic. Read the benchmark report.
What hardware and software were compared?
The test used one Lenovo Yoga 9 15IMH5, so its CPU and GPU shared the laptop’s chassis and cooling system. The CPU was an Intel Core i7-10750H with six cores, twelve threads, and AVX2. The GPU was an NVIDIA GeForce GTX 1650 Ti with Max-Q Design and 4096 MiB of memory.
#1 Best Overall
- Desktop-Level Performance, Anywhere: Get legendary gaming performance with the Intel Core Ultra 9 275HX processor, delivering ultra-smooth gameplay and future-ready AI (Up to 13 NPU TOPS). Offload tasks like background removal and audio optimization to the NPU for seamless streaming and gaming, while Intel Application Optimization enhances performance on classic titles.
- Game-Changing Realism: Powered by NVIDIA Blackwell architecture, GeForce RTX 5070 Ti Laptop GPU unlocks the game changing realism of full ray tracing. Equipped with a massive level of 992 AI TOPS horsepower, the RTX 50 Series enables new experiences and next-level graphics fidelity. Experience cinematic quality visuals at unprecedented speed with fourth-gen RT Cores and breakthrough neural rendering technologies accelerated with fifth-gen Tensor Cores.
- Supreme Speed. Superior Visuals. Powered by AI: DLSS is a revolutionary suite of neural rendering technologies that uses AI to boost FPS, reduce latency, and improve image quality. DLSS 4 brings a new Multi Frame Generation and enhanced Ray Reconstruction and Super Resolution, powered by GeForce RTX 50 Series GPUs and fifth-generation Tensor Cores.
- The Ultimate in Ray Tracing and AI: NVIDIA RTX is the most advanced platform for full ray tracing and neural rendering technologies that are revolutionizing the ways we play and create. Over 700 games and applications use RTX to deliver realistic graphics and incredibly fast performance with cutting-edge AI features like DLSS Multi Frame Generation.
- Immersive Depth and Detail: At 18 inches with a 16:10 aspect ratio, the pristine WQXGA screen offering vibrant colors with up to 100% DCI-P3 operates at a fast 240Hz refresh and 3ms overdrive response time. Alongside the suite of features from NVIDIA G-SYNC and NVIDIA Advanced Optimus, you're guaranteed that whatever's on-screen is a distinct viewing delight.
The machine ran Debian forky/sid with kernel 7.2.6. The reported toolchain was gcc 16.2.0, CUDA 13.4 (V13.4.92), driver 615.71.09, and llama.cpp commit f95b0d9 (build 318). The model was google/gemma-4-E2B-it-qat-q4_0-gguf, described by the author as a 3.35 GB quantization-aware GGUF.
Both devices served the model using llama-server from the same llama.cpp commit. Settings were matched except for -ngl, the GPU-layer-offload setting. Shared settings included an 8192-token context, f16 key/value cache, flash attention, six CPU threads, twelve batch threads, one parallel request, and metrics enabled.
Rank #2
- Intel Core i9 14th Gen 14900HX 1.6GHz Processor, NVIDIA GeForce RTX 5070 8GB GDDR7, 32GB DDR5-5600 RAM
- 1TB PCIe Gen4 x4 NVMe M.2 SSD
- 15.1" WQXGA OLED Glossy Display
- Gigabit LAN, 2x2 WiFi 7 (802.11be), Bluetooth 5.4
- 4.19 lbs. (1.90 kg),Windows 11 Home
How was the benchmark run?
The author used an ABBA sequence: CPU, GPU, GPU, CPU. Before each pass, the script waited at least 120 seconds and until the CPU package temperature was no higher than 50 °C and the GPU temperature no higher than 45 °C. Each pass covered four prompt lengths—94, 516, 998, and 1959 tokens—and two output lengths—32 and 128 tokens—with three repeats per combination. Requests ran one at a time. The report says the prompt cache was cold and the page cache warm.
Alternating the order helps expose whether a device’s result depends on being tested first or after the chassis has warmed. In this run, the ABBA aggregate was 4.14× for decode and 3.42× for prefill. The CPU-first results were 4.09× and 3.38×, respectively; GPU-first results were 4.17× and 3.47×. The author characterized the order effect as about 2% of the ratio on this laptop.
Rank #3
- [Top Performance Processors] KAIGERR Light gaming laptop R7-5700U by ΑΜD ZEN 3 architecture, matched with 16MB of L3 cache, built by TSMC 7nm process with 8 cores & 16 threads (turbo up to 4.3GHz). KAIGERR office light gaming laptops makes it easy to qualify for your PC work and PC games which have amazing loading and processing power for a smoother PC used experience
- [Huge Capacity Storage] KAIGERR laptop comes with 16GB SODIMM DDR4 RAM, advantages of large operating memory capacity both can reduce read latency of memory data and improve CPU utilization. KAIGERR laptop computer configured with an M.2 2280 NVMe 512GB SSD which offers fast startup and loading of applications, as well as a large amount of storage space for your various files
- [Brilliant Display & Integrated Graphics] KAIGERR Light gaming laptop features an innovative thin-bezel display that provides more usable onscreen space for immersive FHD viewing. KAIGERR laptop integrates with ΑΜD Radeon Graphics and delivers strong graphics processing like a rich level of image detail making it possible to play computer games or edit pictures with a great experience on this laptop
- [Rich Interfaces & Wireless Connectivity] KAIGERR traditional laptop offers a variety of connectivity options, including HDMI, Type-C, 3.5mm TRRS Jack, Memory Card Slot and USB3.2 ports. You can easily connect to various devices and peripherals to expand your capabilities. Mini laptop computers equipped with WiFi6 & Bluetooth 5.2 which offer strong wireless signal, fast wireless connections, and reliable transmission speed
- [Portable Design & Durable] KAIGERR laptop compact design makes it easy to carry with you wherever you go. Also, you can enjoy the benefits of a powerful computer without the bulk of a traditional desktop. KAIGERR laptop computers are built with high-quality components and designed to handle heavy workloads and deliver consistent performance and longevity. If you encounter any problems, please contact us and we will help you solve the problem within 12 hours
Did thermal drift affect the result?
The reported repeat behavior differed between devices. GPU pass-to-pass decode drift had a median of +0.6%, ranging from 0.0% to +1.3%. CPU drift had a median of −1.1%, ranging from −9.3% to +0.1%. The CPU’s second pass recorded twice as many throttle events as its first, and CPU repeats within a cell spread by as much as 14.05%; GPU repeats spread by no more than 1.14%. That supports describing the CPU readings as more thermally variable in this run, not claiming that GPUs are inherently more stable.
A temperature gate before each pass cannot by itself show what temperatures or clocks did during that pass. A commenter raised that limitation and suggested tracking per-pass variance alongside clocks and temperatures. The concern is a methodological question, not evidence that a thermal event invalidated the reported ratio.
Rank #4
- 【ENGINEERED FOR SPEED】The KAIGERR 2026 RX16 laptop is equipped with the powerful AMD Ryzen 7 H255 processor (8C/16T, up to 4.9GHz), delivering superior performance and responsiveness. This upgraded hardware ensures a smooth experience, fast loading times, and high-quality visuals. It provides an immersive, lag-free experience. Its performance is far moere than 30% better than AMD R7 5700U/5800U/5825U/6600HX/7735HS.
- 【Advanced Dual-Fan Cooling】KAIGERR’s dual-fan system expels heat faster than standard designs, drastically reducing thermal buildup during intense gaming or work. Optimized airflow keeps components cool, prevents throttling, and maintains smooth, sustained performance—all while staying quiet. Stay cool, play longer.
- 【INSPIRE YOUR POSSIBILITIES】 The laptop on sale comes with 16GB DDR5 memory and a 512GB M.2 NVMe SSD for faster response times and ample storage. Dual-channel DDR5 memory supports upgrades to 64GB (2x32GB), and NVMe/NGFF SSD can be upgraded to 4TB, providing plenty of space for all your favorite videos/files.
- 【Vivid 16.0" IPS Display】Featuring a wide color gamut and high refresh rate, the 16.1" IPS screen delivers smoother motion, richer colors, and exceptional detail—surpassing standard displays in both accuracy and immersion. Whether gaming, streaming, or creating, every frame appears lifelike and dynamic for a truly engaging visual experience.
- 【KAIGERR: Quality Laptops, Exceptional Support.】Enjoy peace of mind with unlimited technical support and 12 months of repair for all customers, with our team always ready to help. If you have any questions or concerns, feel free to reach out to us—we’re here to help.
What does this mean for someone serving Gemma locally?
For interactive use on this particular machine and configuration, the GPU was the faster option. The CPU remains a usable fallback when the GPU is unavailable or occupied, though its measured decode rate was substantially lower in this test. The result does not establish which device is preferable for another laptop, a different model size or quantization, multiple simultaneous requests, or another software build.
Recommended Free Tools
When comparing another benchmark with this one, check whether it matches the factors that can change the outcome:
- Model, quantization, and model size.
- Exact CPU and GPU, including whether they share a laptop cooling system.
- Software commit, build options, and thread settings.
- Prompt and output lengths, along with concurrency.
- Run order and temperature controls.
- Whether it reports decode, prefill, end-to-end latency, and repeat variation.
How strong is the evidence?
This is a single-machine benchmark covering one GGUF model, one llama.cpp commit, one concurrency setting, and eight paired prompt/output combinations. The author mentions an earlier run with a 4.27× decode lead and 3.63× prefill lead, but the commit, thread flags, and run order all changed together. The difference cannot be attributed to any single change, so the two runs are not a controlled before-and-after comparison.
The report links run details, but no independent replication or audit was verified. The author also said replicating the test on a second machine would be interesting but difficult because this laptop is uncommon. The measurements are useful as a carefully specified result for this setup; they do not establish a general GPU-versus-CPU performance ratio.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.
Free tools Windows power users keep installed
One-click scans. No signup required.

