Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

One slow GPU rarely stalls a training run on its own. In synchronous distributed training, every worker must reach the same synchronization point before any of them can move on, so the whole job advances at the pace of its slowest participant. The harder question is why that participant is late, and the documented causes are more varied than the phrase “slow GPU” suggests.

Why one late worker holds up the whole job

Data parallelism: the gradient exchange

In data parallelism, each worker processes its slice of a global batch, and the workers then exchange gradients before the optimizer step. In practice that exchange is a barrier. A worker that arrives late forces every faster worker to idle until it catches up. ZeRO and FSDP change which model state is sharded and which collectives run, typically reduce-scatter and all-gather rather than a single all-reduce, but they still contain a coordination point at each step, so a late participant stalls them in the same way.

Pipeline parallelism: delay becomes a bubble

Pipeline parallelism splits the model’s layers into stages that pass microbatches along. If one stage carries more work than its neighbors, or a microbatch takes longer than expected, the stages downstream sit idle. Those idle gaps are pipeline bubbles. A delayed microbatch can push later work back within the same step, and the effect can spread beyond the step in which it began.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Tensor and context parallelism: stalls inside a group

Tensor and context parallelism split a single layer or sequence across a small group of GPUs that must combine partial results. A device that falls behind holds back its group peers at each exchange. The effect stays within the group, but because these exchanges are frequent, a small delay can recur many times in one step.

#1 Best Overall
ASUS Dual Radeon RX 9060 XT 16GB GDDR6 Gaming Graphics Card
  • Axial-tech fans now feature a smaller fan hub that facilitates longer blades and a barrier ring that increases downward air pressure
  • 2.5-slot design allows for greater build compatibility while maintaining cooling performance
  • 0dB technology lets you enjoy light gaming in relative silence
  • Dual BIOS switch lets you toggle between Quiet and Performance BIOS profiles
  • Dual ball fan bearings last up to twice as long as sleeve bearing designs

A straggler is a late arrival, not necessarily a broken card

The useful distinction is between a process that finishes its own work late and processes that finish early and then wait. The waiting processes show the longest synchronization time in a profile, which is why they are easy to mistake for the problem. The straggler is the process whose work ended last.

Hardware can be that cause, but the documented cases name several other sources as well, covered below. Diagnosis therefore starts with what the late process was doing, not with what it runs on.

Rank #2
GIGABYTE GeForce RTX 5070 Ti Gaming OC 16G Graphics Card, 16GB 256-bit GDDR7, PCIe 5.0, WINDFORCE Cooling System, GV-N507TGAMING OC-16GD Video Card
  • Powered by the NVIDIA Blackwell architecture and DLSS 4
  • Powered by GeForce RTX 5070 Ti
  • Integrated with 16GB GDDR7 256bit memory interface
  • PCIe 5.0
  • WINDFORCE cooling system

Where stragglers come from

Uneven work across stages and microbatches

  • Unequal pipeline stages. The OSDI ’25 study of ByteDance’s LLM training cluster found that work imbalance between pipeline stages caused many observed stragglers. If layers or operations are distributed unevenly, the heaviest stage becomes the bottleneck and the others wait on it.
  • Sequence-length imbalance. The same study found that sequence-length imbalance between microbatches was a major cause. Microbatches with longer sequences require more computation, so the rank that holds them can finish late even when every GPU is identical.

Input data and host-side pauses

  • Outlier-sized examples that make one batch much heavier than its peers, as PyTorch’s engineering discussion of DDP describes.
  • Slow or unstable data loading, including unstable network I/O during data transfer.
  • Variable transformation cost for transforms applied on the fly to each example.
  • Garbage-collector pauses. The ByteDance study identified these as a cause in its training cluster. A pause keeps that process from reaching the next synchronization point until it finishes.

Communication paths

  • Network congestion on the path a collective uses.
  • RNIC or switch defects, and topology asymmetry. The NSDI ’26 PIPEMORPH work cites these as communication-straggler conditions in pipeline training.

Device interruptions

Some newer work addresses devices that drop out or become unavailable during a run. This is related to a persistently slow worker but is a different problem: the group has to reshape its parallelism around the devices it still has. NVIDIA’s 2026 technical blog describes one such approach, which it calls NTP, and labels it experimental.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What the ByteDance trace shows, and where its reach ends

The OSDI ’25 paper Understanding Stragglers in Large Model Training Using What-if Analysis analyzed a five-month trace from ByteDance’s LLM training cluster, covering January through May 2024. Within that trace:

Rank #3
Sale
GIGABYTE GeForce RTX 5060 WINDFORCE OC 8G Graphics Card, Cooling System, 8GB 128-bit GDDR7, PCIe 5.0, Manufactured by NVIDIA, DisplayPort & HDMI - Video Output Interface, GV-N5060WF2OC-8GD Video Card
  • Powered by the NVIDIA Blackwell architecture and DLSS 4
  • Powered by GeForce RTX 5060
  • Integrated with 8GB GDDR7 128bit memory interface
  • PCIe 5.0
  • WINDFORCE cooling system
  • 42.5% of jobs were at least 10% slower due to stragglers.
  • For the jobs at the tail of the distribution, the study estimates that stragglers could waste up to 45% of allocated resources.

These figures describe one cluster over one window, not a general straggler rate for other clusters, models, or schedulers. The paper also reports that most steps in a straggling job showed similar slowdowns, which the authors read as evidence of persistent problems rather than transient environmental noise:

“Most steps incur similar slowdowns within a straggling job, suggesting that they are often not caused by transient environmental issues but are rather caused by persistent problems.”

Rank #4
Sale
GIGABYTE Radeon RX 9070 XT Gaming OC 16G Graphics Card, PCIe 5.0, 16GB GDDR6, GV-R9070XTGAMING OC-16GD Video Card
  • Powered by Radeon RX 9070 XT
  • WINDFORCE Cooling System
  • Hawk Fan
  • Server-grade Thermal Conductive Gel
  • RGB Lighting

In the same trace, computation operations were slower more often than communication operations, and the study found no positive correlation between job size and straggler-related slowdown. Its method models the job’s operation dependencies and simulates the effect of removing straggler time, rather than labeling every slow step as a hardware fault. That is the key difference from a per-step timer: the question it answers is how much faster the job would run if the straggling were removed.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How to find the straggler before changing anything

Start with per-rank timelines, not an aggregate step time. Aggregate time tells you the job is slow; per-rank timelines tell you which rank’s work ended last. In this article, a “rank” means one training process, usually one GPU.

Best Value
Sale
ASUS Prime Radeon RX 9070 XT 16GB GDDR6 OC Edition Gaming Graphics Card
  • Axial-tech fans now feature a smaller fan hub that facilitates longer blades and a barrier ring that increases downward air pressure
  • Phase-change GPU thermal pad helps ensure optimal heat transfer, lowering GPU temperatures for enhanced performance and reliability
  • 2.5-slot design allows for greater build compatibility while maintaining cooling performance
  • Dual-ball fan bearings last up to twice as long as standard conventional sleeve bearings designs
  • 0dB technology lets you enjoy light gaming in relative silence

Read the synchronization interval correctly

A rank that finishes its compute early shows a long all-reduce interval, because it spends that time waiting for the others. PyTorch’s DDP example makes this point: the process with the highest reported synchronization cost is often one of the fast processes, not the straggler. Compare the same step across ranks. The straggler is the one with the shortest wait and the latest end of its pre-collective work.

Diagnostic checklist

  1. Choose a slow step and export a per-rank profile for that step. A profiler trace such as one from torch.profiler records the timeline for each process.
  2. For each rank, record two numbers: the time from step start to the first synchronizing collective, and the duration of that collective.
  3. Identify the rank whose pre-collective work ends last. That rank is the straggler candidate; long collectives on the other ranks are the symptom.
  4. Check that rank’s inputs: sequence lengths and example sizes in its microbatches, and time spent waiting on the data loader.
  5. Check its host side for garbage-collection pauses and for CPU gaps during which the GPU is idle.
  6. For pipeline jobs, compare busy time across stages to locate the heaviest stage and the bubbles it creates.
  7. Only then examine the network. If waits concentrate on one collective or one peer pair, set NCCL_DEBUG=INFO to log which communicator and transport are in use.

The OSDI ’25 paper reports that parts of its analysis were incorporated into SMon, a tool deployed in the ByteDance cluster and used by its on-call team to detect and address stragglers. That is a documented internal operational example; the paper does not present SMon as a generally available product.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Mitigations and what each one costs

Mitigations differ in the cause they address and in what they change about synchronization. The table compares the main options.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Approach Cause it addresses Mechanism Trade-off or limit Evidence context
Rebalance the work Uneven pipeline stages, sequence-length imbalance, data loading Move layers or operations between stages; manage sequence-length spread across microbatches; fix loader and transform costs. Requires identifying the delayed operation; no single change covers every cause. Causes documented in the ByteDance trace (OSDI ’25) and PyTorch’s DDP discussion.
Hierarchical SGD Random slow processes Synchronize often within small groups and less often across groups, limiting how far one process’s delay reaches. Changes synchronization cadence; warmup and hierarchy settings affect convergence and model parity. PyTorch describes an implementation and illustrative experiments.
Asynchronous SGD Waiting at every synchronous boundary Workers update without waiting at each synchronous boundary. Stale gradients; see the note below the table. A 2018 AISTATS paper analyzes the runtime/error trade-off. IBM’s description of grouped synchronization presents it as an intermediate balance.
Pipeline rescheduling and communication offload (PIPEMORPH) Communication delays that create pipeline bubbles Adapt scheduling around communication delays; move communication operations to host memory and CPU-side RDMA to reduce GPU head-of-line blocking. Specialized systems work; results are experimental and setting-specific. The NSDI ’26 PIPEMORPH paper reports 1.2–3.5× iteration-time improvement in its tested settings.
Adaptive tensor parallelism (NVIDIA’s NTP) Device unavailability during a run Reconfigure a replica to use the GPUs that remain, overlapping resharding with computation and synchronization. Experimental; outcomes depend on hardware, power, and software assumptions. NVIDIA’s 2026 technical blog, which labels the approach forward-looking and experimental.

The 2018 AISTATS paper by Sanghamitra Dutta, Gauri Joshi, Soumyadip Ghosh, Parijat Dube and Priya Nagpurkar states the central trade-off of asynchronous training: “Asynchronous methods can alleviate stragglers, but cause gradient staleness that can adversely affect convergence.”

Quick Recap

Bestseller No. 1
ASUS Dual Radeon RX 9060 XT 16GB GDDR6 Gaming Graphics Card
ASUS Dual Radeon RX 9060 XT 16GB GDDR6 Gaming Graphics Card
0dB technology lets you enjoy light gaming in relative silence; Dual BIOS switch lets you toggle between Quiet and Performance BIOS profiles
$529.00
Bestseller No. 2
GIGABYTE GeForce RTX 5070 Ti Gaming OC 16G Graphics Card, 16GB 256-bit GDDR7, PCIe 5.0, WINDFORCE Cooling System, GV-N507TGAMING OC-16GD Video Card
GIGABYTE GeForce RTX 5070 Ti Gaming OC 16G Graphics Card, 16GB 256-bit GDDR7, PCIe 5.0, WINDFORCE Cooling System, GV-N507TGAMING OC-16GD Video Card
Powered by the NVIDIA Blackwell architecture and DLSS 4; Powered by GeForce RTX 5070 Ti; Integrated with 16GB GDDR7 256bit memory interface
$1,162.49
SaleBestseller No. 3
GIGABYTE GeForce RTX 5060 WINDFORCE OC 8G Graphics Card, Cooling System, 8GB 128-bit GDDR7, PCIe 5.0, Manufactured by NVIDIA, DisplayPort & HDMI - Video Output Interface, GV-N5060WF2OC-8GD Video Card
GIGABYTE GeForce RTX 5060 WINDFORCE OC 8G Graphics Card, Cooling System, 8GB 128-bit GDDR7, PCIe 5.0, Manufactured by NVIDIA, DisplayPort & HDMI - Video Output Interface, GV-N5060WF2OC-8GD Video Card
Powered by the NVIDIA Blackwell architecture and DLSS 4; Powered by GeForce RTX 5060; Integrated with 8GB GDDR7 128bit memory interface
$459.99
SaleBestseller No. 4
GIGABYTE Radeon RX 9070 XT Gaming OC 16G Graphics Card, PCIe 5.0, 16GB GDDR6, GV-R9070XTGAMING OC-16GD Video Card
GIGABYTE Radeon RX 9070 XT Gaming OC 16G Graphics Card, PCIe 5.0, 16GB GDDR6, GV-R9070XTGAMING OC-16GD Video Card
Powered by Radeon RX 9070 XT; WINDFORCE Cooling System; Hawk Fan; Server-grade Thermal Conductive Gel
$814.99
SaleBestseller No. 5
ASUS Prime Radeon RX 9070 XT 16GB GDDR6 OC Edition Gaming Graphics Card
ASUS Prime Radeon RX 9070 XT 16GB GDDR6 OC Edition Gaming Graphics Card
0dB technology lets you enjoy light gaming in relative silence; Dual BIOS switch lets you toggle between Quiet and Performance BIOS profiles
$829.00

Choosing among them

  1. If profiles show that one rank’s pre-collective work ends last because of its data or sequence lengths, fix the workload first.
  2. If waits concentrate on one communication path, examine congestion and topology before changing the training algorithm.
  3. If a device drops out during a run, reconfiguring parallelism is the relevant class of fix; check the maturity column above before depending on it.
  4. Change synchronization semantics, through hierarchical or asynchronous methods, only after confirming the change does not hurt convergence relative to the synchronous baseline.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.