Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

On my dual Tesla P40 system, row splitting once delivered substantially higher throughput than layer splitting. When a later model and CUDA setup stopped working with that mode, the practical lesson was not that row split had vanished for everyone: it was that the performance rule I had tuned around applied only to a particular build, model, backend, and workload. I had to test the available modes again—and recover throughput using other features.

What row splitting did for my dual Tesla P40 setup

In my earlier setup, I measured about 12–14 tokens per second with row splitting, compared with about 7 tokens per second with layer splitting. In an earlier 72B-model configuration, I reported approximately 10.3 generated tokens per second and 60 prompt tokens per second with the model fully resident on the GPUs and row split enabled. Those figures are my measurements on my own hardware and workloads, not portable expectations or independently replicated benchmarks. The original account

The key point is that a split-mode result is conditional. The model, GPU backend, binary, device arrangement, and workload all matter. A mode that is faster in one combination can behave differently—or fail—in another.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Why I had to test the configuration again

Changing several factors hid the regression

In a March comparison, I changed multiple factors at once, which obscured a significant prompt-processing regression. A later one-variable-at-a-time comparison made the differences clearer: row split worked on the original binary, layer split ran at about half the speed, and graph split crashed on Pascal with an illegal-memory-access error. These were outcomes in my particular tests, not rules for every llama.cpp release or Pascal system. My account of the comparisons

#1 Best Overall
ASUS ROG G700 (2025) Gaming Desktop PC, Intel® Core™ Ultra 7 265F Processor, NVIDIA® GeForce RTX™ 5070, 1TB M.2 NVMe™ PCIe® 4 SSD, 16GB DDR5 RAM, Windows 11 Home, G700TF-DS774
  • Fearless ROG Design – The G700’s dual-glass chassis showcases iconic ROG design with the ROG Slash and Aura Sync RGB lighting. Its 58L capacity supports triple-slot GPUs.
  • Unstoppable Power – Equipped with the Intel Core Ultra 7 265F processor, NVIDIA GeForce RTX 5070 GPU, 16GB DDR5 RAM, and 1TB SSD PCIe 4.0 storage for seamless gaming and multitasking.
  • Optimized Thermals – Stay cool with a quad-fan system, while dust filters and efficient airflow ensure long-term reliability.
  • Advanced Connectivity – Game without lag with 2.5Gbps Ethernet, Wi-Fi 6, and versatile ports. Dolby Atmos audio and AI noise cancellation enhance sound and communication.
  • Ready for Upgrades – Designed with tool-less access, easily swap out components, ensuring future-proof performance for years to come.

The model and backend combination changed the result

In my multi-GPU CUDA setup, I found that Gemma 4’s shared KV layers, represented as tensor views, caused row split to fail. My Qwen stacks continued to use row split. This is my architecture- and configuration-specific account; it should not be read as proof that Gemma 4 or row split universally fails across all builds and backends. My account of the model behavior

A separate upstream report filed July 12, 2026 documents a row-split failure on one CUDA build in a mixed CUDA/ROCm setup. That report is evidence of a particular compatibility problem, not evidence of universal removal. It also describes other split-mode failures in that environment, underlining why the backend and device mix need to be part of any comparison. The July 12, 2026 issue

Rank #2
Sale
CyberPowerPC Gaming PC, AMD Ryzen 5 5500, Radeon RX 6500 XT 4GB
  • System: AMD Ryzen 5 5500 3.6GHz 6 Cores | AMD B550 Chipset | 8GB DDR4 | 500GB PCIe 4.0 NVMe SSD | Windows 11 Home
  • Graphics: AMD Radeon RX 6500 XT 4GB Graphics | 1x HDMI | 1x DisplayPort
  • Connectivity: 4 x USB-A 3.2 | 4 x USB-A 2.0 | 1 x LAN | WiFi 5 | Bluetooth 5.0 | 7.1 Channel Audio
  • Tempered Side Case Panel | Custom RGB Lighting | Keyboard and Mouse
  • 1 Year Parts & Labor Warranty, Free Lifetime Tech Support

Was row split deleted from llama.cpp?

Not universally, based on the documentation available for this account. The llama.cpp server README retrieved around October 7, 2026 lists -sm, --split-mode {none,layer,row,tensor}. It describes layer as the default, with layers and KV split across GPUs; row splits weights by rows; and tensor mode is experimental. The same README lists --parallel / -np for the number of parallel sequences to decode. Server README

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The CLI README retrieved around October 7, 2026 likewise lists row split, parallel sequences, and speculative decoding modes including draft-mtp. These are mutable master documentation pages rather than pinned release notes, so the availability of a flag in a particular release or binary must be checked against that version. CLI README

Rank #3
Sale
WIWB Gaming PC Desktop, GeForce RTX 3050 8GB GDDR6, AMD Ryzen 7 4700LE
  • 8-Core 16-Thread Processing Power – Powered by the Ryzen 7 4700LE processor with Zen 2 architecture, delivering 8 cores and 16 threads with a boost clock up to 4.2GHz. Effortlessly handle multitasking, streaming, content creation, and demanding applications simultaneously without slowdowns.
  • GeForce RTX 3050 8GB Graphics – Equipped with 8GB GDDR6 dedicated VRAM and real-time ray tracing support. Experience smooth 1080p gaming at 55-60 FPS in AAA titles like Cyberpunk 2077, 70+ FPS in Fortnite, and 90-100 FPS in Apex Legends with DLSS enabled. The 8GB buffer handles modern game textures comfortably – a step above 6GB variants
  • High-Speed Memory & Storage – Paired with 16GB of DDR4 3200MHz dual-channel RAM (16GB), the PC ensures responsive multitasking—whether streaming while gaming or editing videos. It also includes a 512 GB NVMe M.2 SSD for lightning-fast boot times, quick game loads, and ample storage for your game library, creative projects, and files.
  • Next-Gen WiFi 6 Connectivity – Stay connected with the latest WiFi 6 technology for faster speeds, lower latency, and improved network efficiency. Whether you're gaming online, streaming 4K content, or joining video conferences, enjoy stable, high-speed wireless connectivity.
  • Ready-to-Use Value Desktop – Pre-built and ready to go right out of the box. Perfect for gamers, students, content creators, and home office users seeking reliable performance without the hassle of building a PC themselves. The mature AM4 platform with DDR4 memory offers excellent value and proven stability.

My experience was that row split stopped being usable in a particular model/backend combination. That is different from saying upstream removed the mode everywhere. When diagnosing a missing or failing flag, identify the exact build and backend first, then distinguish a removed option from an option that exists but does not work with the current configuration.

How I recovered throughput without a substitute split flag

Parallel slots improved aggregate throughput

In a later stack, I measured 8.46 tokens per second for one stream using layer split. With four parallel slots, my reported aggregate throughput rose to 15.0 tokens per second; at two slots, I reported 12.8 tokens per second. These numbers describe my tests, not single-request latency or a general scaling promise. My later measurements

Rank #4
msi Codex Z2 Gaming Desktop, AMD R7-8700F, RTX 5070, 32GB DDR5, 2TB SSD
  • POWERHOUSE 8-CORE GAMING PERFORMANCE — Driven by the AMD Ryzen 7 8700F with 8 cores and 16 threads, boosting up to 5.0 GHz for smooth, responsive gameplay and the ability to handle AAA titles, streaming, and background tasks all at once
  • NEXT-GEN BLACKWELL ARCHITECTURE — The NVIDIA GeForce RTX 5070 is powered by NVIDIA's cutting-edge Blackwell GPU architecture, delivering a massive generational leap in rasterization and ray tracing performance so you can experience your games the way they were meant to be played.
  • Simplistic Design: Enjoy the latest generation of Windows 11 Home for your everyday needs. *MSI recommends Windows 11 Pro for business use.
  • Cool While Gaming: In conjunction with an ARGB fan Air Cooler, the Codex R2 features four system cooling fans; three in the front and one in the rear to pull in cool air and push heat out of the PC.
  • Turn on the Bright Lights: With the built-in RGB lighting, take your gaming experience to the next level by pressing the MSI LED button to cycle through lighting options. Customize lighting even further with MSI Center software.

Parallelism is useful when the goal is to serve multiple sequences and raise aggregate throughput. It does not mean an individual response necessarily generates faster: a one-stream result and a multi-slot aggregate result measure different things. The server documentation describes --parallel as the number of parallel sequences to decode. Server README

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

MTP speculative decoding improved my single-stream result

I also tested MTP speculative decoding. In my setup, the single-stream result increased from 8.46 to about 13.3 tokens per second, which I reported as a 57% gain. I observed acceptance rates from 0.38 to 0.63 and checked output correctness. These are my reported results, not an independently verified benchmark or a guarantee that MTP will help every model and workload. My MTP measurements

Best Value
KOTIN Prebuilt Gaming PC RTX 5070 12GB, Ryzen 7 9700X, 32GB DDR5, 1TB SSD
  • POWERED BY RTX 5070 12GB + RYZEN 7 9700X - The GeForce RTX 5070 12GB GDDR7 graphics card pairs with an 8-core AMD Ryzen 7 9700X processor to drive smooth 1440p and 4K gameplay, giving this gaming PC the headroom for modern titles, streaming, and creative work.
  • 32GB DDR5 6000MHz MEMORY & 1TB NVMe SSD - 32GB of high-speed DDR5 memory and a 1TB PCIe 4.0 NVMe solid state drive deliver quick load times, smooth multitasking, and generous storage, keeping this prebuilt gaming desktop responsive under heavy workloads.
  • BUILT-IN 11.3-INCH Smart DISPLAY - An integrated smart screen shows real-time CPU and GPU temperatures, usage, and weather while you play, adding a distinctive and functional touch to your battlestation.
  • 850W 80+ GOLD POWER SUPPLY, 360MM LIQUID COOLING & WiFi 7 - An 850W 80 Plus Gold certified power supply provides stable, efficient power with headroom for future upgrades, while a 360mm AIO liquid cooler, WiFi 7, and an ARGB mid-tower case keep the Ryzen 7 CPU cool and connected in a clean build.
  • READY TO PLAY OUT OF THE BOX - Arrives fully assembled and tested with Windows 11 Home pre-installed, so your prebuilt gaming computer is ready to set up in minutes. Assembled in the USA, and backed by a one-year limited warranty and lifetime free technical support.

Those changes were ways to recover throughput in the later stack; neither was a direct replacement split-mode flag. The CLI README lists draft-mtp among speculative decoding modes, but exact support and behavior depend on the binary and configuration being used. CLI README

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

A practical way to re-measure after a flag or model change

  1. Record the baseline. Note the llama.cpp binary or commit, build options, backend, GPU models and arrangement, model and quantization, split mode, workload, and whether you are measuring prompt processing, generation, or concurrent aggregate throughput.
  2. Confirm the mode exists in your build. Check the documentation or help output for the exact release and backend rather than assuming that a mutable upstream page describes an older binary.
  3. Change one factor at a time. Hold model, prompt, generation settings, and workload constant while comparing modes. If a mode fails, record the exact error and configuration instead of treating the result as a performance score.
  4. Separate latency from throughput. Measure one stream for single-request behavior. If you use parallel slots, report the slot count and aggregate throughput separately.
  5. Verify output as well as speed. For speculative decoding or other execution changes, check correctness alongside generation rate.
  6. Repeat after meaningful changes. A new model architecture, backend, binary, or device mix can invalidate a previously useful tuning rule.

Current documentation describes the available modes but does not establish a universal performance ranking. Choose based on support in the specific release and backend, model compatibility and correctness, single-request latency, aggregate throughput under concurrency, and stability under the workload you actually run. My results and configuration limits Server mode descriptions CLI mode descriptions

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.