Federated learning can run slowly or produce weaker results on some devices because clients differ in compute capacity, usable memory, network conditions, availability, and local data. Start by identifying where the problem occurs—loading, local training, communication, participation, update freshness, or global model quality—then measure that stage separately. A faster round is not necessarily a better training run if it repeatedly excludes clients or data groups.
What does “poor performance” mean in federated learning?
The phrase can describe several different failures, and each points to a different investigation. A client may be unable to load or train the model, take unusually long to finish local work, fail to send or receive data, miss a round, or submit an update too late to be useful. Separately, the global model may perform poorly because participating clients do not adequately represent the data the system needs to learn from.
- Client feasibility: the device cannot load the model or complete its assigned training work.
- Client speed: training completes, but takes longer than it does on comparable clients.
- Communication: model downloads or update uploads are delayed or fail.
- Participation or freshness: the device is unavailable, misses selection or deadlines, or contributes an update that is old when aggregated.
- Global quality: the resulting model is weak, potentially because of data coverage or distribution rather than device speed.
These symptoms can overlap. Record which stage is slow or failing before changing the training configuration.
Why do some devices struggle?
Compute and memory capacity vary
Clients can differ in processor performance, available memory, accelerators, software generation, power conditions, and competing workloads. The useful distinction is between a hard constraint and a soft constraint: a hard constraint prevents a client from running the workload, while a soft constraint allows it to run more slowly, potentially causing it to miss a deadline or become a straggler. Memory pressure can prevent participation even when the model’s parameter count appears manageable, because training activations also use memory.
#1 Best Overall
A 2023 survey by Pfeiffer, Rapp, Khalili, and Henkel reports smartphone computation ranging from 1010 to 1012 FLOPS and memory from 512 MB to 8 GB as an illustration of device variation; these are not specifications for current phone models or a direct predictor of training time. The survey also discusses a literature example in which one smartphone has about one hundredth the peak performance and one eighth the memory of a high-end smartphone. Those figures illustrate a particular comparison, not a universal ratio. See the ACM Computing Surveys article, published 17 July 2023.
Communication and device availability affect deadlines
A client can finish its local training but still be late because of low throughput, high latency, or an unreliable connection. Downloading a model and uploading an update should therefore be measured separately from local computation. Communication and computation can also compete for a device’s energy and other resources. A client that is not available when a round runs contributes nothing, regardless of its theoretical compute capacity. The NeurIPS 2023 paper FLuID: Mitigating Stragglers in Federated Learning addresses stragglers in this broader setting.
Rank #2
Local data amounts and distributions differ
Clients may hold different amounts of local data. If a client’s work is defined by examples or batches, that difference can affect its training time. Data can also be non-IID: clients may have different data distributions rather than representative samples of one common distribution. Consequently, a global-quality problem is not proof that slower hardware is at fault; it may instead reflect which clients participated and which data groups were represented.
How should you troubleshoot a slow or unreliable client?
Use the sequence below to narrow down the failing stage. It is a practical diagnostic approach, not a standardized protocol with universal pass/fail thresholds. The reviewed sources do not establish general acceptable limits for latency, memory headroom, bandwidth, or update age.
Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Repair Windows errors before they cause bigger problemsFix Now →- Name the symptom and stage. Identify whether the client fails to load, trains slowly, has communication delays or failures, misses rounds, submits stale updates, or is associated with weak global quality. Record when it happens and whether it affects one client class or changes over time.
- Check whether the workload fits. Look for model-loading failures, training failures, and memory pressure. Distinguish a client that cannot run the workload from one that completes it slowly. Consider training activations as well as model parameters when assessing memory use.
- Measure local compute and contention. Compare local training duration for the same workload across comparable clients and across rounds. Check whether concurrent applications or changes in device state coincide with slowdowns. A device’s own history can help identify variation, but a correlation does not establish a single cause.
- Split communication from computation. Record model-download delay or failure, local training duration, and update-upload delay or failure as separate measurements. This reveals whether work finished locally but could not be delivered on time.
- Trace participation and update age. For each round, distinguish eligibility, availability, selection, completed work, and successful submission. At aggregation, note whether updates are current. In asynchronous systems, also inspect whether slow clients’ updates are older and whether faster devices contribute disproportionately often.
- Compare data quantity and representation. Record examples or local steps per client, then check whether dropped or deprioritized clients share a distinct data distribution. This helps separate a resource issue from a coverage issue.
- Change one system choice at a time and track multiple outcomes. Monitor round time, participation and data coverage, convergence, and final model quality; include energy or resource use when your deployment measures it. A shorter round by itself does not establish an overall improvement.
Which mitigation should you try?
Choose based on the measured bottleneck, and assess speed alongside model quality, client participation, representation, communication burden, and energy or resource use where measured. The sources do not establish one best setting for every workload.
| Option | When it may help | Trade-off to measure |
|---|---|---|
| Resource-aware client selection | When clients’ compute or communication limits are causing stragglers or long waits. | Repeatedly excluding resource-poor clients can reduce data coverage or harm results if device resources are associated with non-IID data distributions. Track which clients and data groups are omitted. ACM survey (2023). |
| Adjust workload to client capability | When clients can train but struggle with the work demanded of them. Heterogeneity-aware approaches can vary resources or work across clients. | There is no universal local-epoch count, batch count, or model size established by these sources. Check feasibility and straggling against coverage and model quality. ACM survey (2023). |
| Asynchronous or partially asynchronous aggregation | When waiting for every slow client is delaying progress. | Updates may be stale, and faster clients may contribute more often, with potential convergence or accuracy costs. The CVPR Workshops 2023 paper TimelyFL reports drawbacks for an asynchronous baseline in its evaluated scenarios; that finding does not establish the result for every asynchronous design. |
| Reduce communication burden | When transfers are a measured bottleneck. The survey discusses reducing model structure to lower computation and communication burden, as well as compression and quantization methods for communication. | Verify any model-quality effects; the sources do not identify one universally appropriate compression or quantization setting. ACM survey (2023). |
| Benchmark device and state variation | When you need to evaluate differences in devices and their changing states, rather than data heterogeneity alone. | Benchmarking provides evaluation context; it does not by itself remove a production bottleneck. The CVPR 2024 paper FLHetBench focuses on device and state heterogeneity. |
How do synchronous and asynchronous rounds change the failure mode?
Synchronous aggregation
A synchronous round waits for selected clients to return their updates or for a deadline or participation rule to take effect. A slow client can therefore delay aggregation. Pfeiffer and co-authors describe this straggler effect in their 2023 survey: “If a device k in the set C^t takes longer than others, then it delays the synchronous aggregation and, hence, slows down the overall FL training.” Whether a particular client actually blocks a round depends on the system’s aggregation and deadline policy.
Asynchronous aggregation
An asynchronous design can make progress without waiting for every slow client, but it changes the risk: an update may be stale by the time it is applied, while fast clients may contribute more frequently. Measure update age and contribution frequency as well as elapsed time; otherwise a faster system can conceal uneven participation or a convergence-quality cost. The TimelyFL paper evaluates one asynchronous baseline, not every possible asynchronous method.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.
Recommended Free Tools

