Free tools Windows power users keep installed
One-click scans. No signup required.
The learning rate sets how far a neural network’s parameters move on each optimization step. A rate that is too small can make training painfully slow; one that is too large can cause loss to oscillate or diverge. The best choice depends on the optimizer, model, data, batch size, and training stage—there is no universal value that guarantees the best accuracy.
What the learning rate changes
During training, an optimizer uses gradients to update the model’s parameters. The learning rate scales that update: larger values take longer steps through the loss surface, while smaller values take shorter ones. It therefore affects both how quickly training progresses and whether those steps remain stable.
A larger rate can reduce the number of updates needed to make early progress when the steps are well-sized. But the loss surface is not equally steep in every direction. In classical stability analysis, the largest eigenvalue of the loss Hessian—the measure of curvature in its sharpest direction—helps determine the maximum stable step size. A rate that exceeds what the current curvature can tolerate may overshoot a minimum instead of moving toward it.
What happens when the rate is too low or too high?
Too low: stable but slow progress
Small updates are generally more controlled, but they can require many more training steps to reduce loss. If training appears stable yet makes little progress, the rate may be too low—or another part of the setup may be limiting learning. Check the loss trend over a meaningful number of updates rather than judging from a single step.
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Fix the driver behind crashes, sound loss and screen glitches3Repair Windows errors before they cause bigger problems#1 Best Overall
Too high: overshooting, oscillation, or divergence
If updates are too large for the local curvature, the optimizer can repeatedly jump across a low-loss region. Training loss may oscillate instead of settling, or grow until training becomes unstable. Lowering the rate is a direct way to test whether step size is the cause, though gradients, data, and other settings can also produce unstable training.
Near the stability boundary: loss need not fall smoothly
Training does not always follow a simple pattern in which loss decreases monotonically at every step. Recent work describes an “edge of stability” regime where loss decreases non-monotonically while sharpness stays near the stability boundary. Galli and colleagues’ ICML 2026 paper reports that the product of step size and sharpness can remain above the edge-of-stability threshold of 2 throughout training; this is a finding in their study, not a universal threshold to apply to every model and optimizer. Read the ICML 2026 study.
Rank #2
How learning rate affects convergence and accuracy
There is no single accuracy gain or percentage improvement that transfers across architectures and tasks. A higher rate may reach a target quality in fewer updates if it stays stable; a rate that is too high can prevent useful convergence. Compare rates by how quickly they lower training loss, how many updates or how much time they need to reach a target validation metric, how stable training remains, and the compute required.
In a 2003 study covering a 20,000-instance speech-recognition task and 26 other learning tasks, Wilson and Martinez found that online training could safely use a larger learning rate than batch training and converge in fewer passes through the data, with no apparent accuracy difference on the tested tasks. They attributed the result to online training’s ability to follow curves in the error surface during an epoch. That task-specific result does not establish that online training or larger rates will preserve accuracy in every modern neural network. Read Wilson and Martinez’s 2003 paper.
Recommended Free Tools
Rank #3
Why a larger rate does not guarantee better generalization
Generalization means performance on data the model did not train on. In some settings, larger learning rates are associated with flatter solutions or useful implicit regularization, and minibatch noise can also contribute to generalization behavior. These effects depend on the training setup; a larger rate is not a general-purpose fix for overfitting.
Galli and colleagues’ ICML 2026 experiments report that reaching globally flat regions too early can slow convergence and hurt generalization in their settings. This illustrates an important qualification: even when flatness is useful, when and how the optimizer reaches a region can matter. Smith, Elsen, and De study the role of minibatch noise in generalization; their work likewise does not imply a universal rule that increasing the learning rate improves validation performance. Read Smith, Elsen, and De’s ICML 2020 paper.
Rank #4
How batch size and learning rate interact
Batch size determines how many examples contribute to a gradient estimate for each update. Changing it changes the training dynamics, so a learning rate that worked at one batch size may not work at another. NeurIPS 2019 research provides theoretical and empirical evidence that the batch-size-to-learning-rate ratio should not be too large for good generalization. It supports tuning the two together, not a universal conversion formula that produces an optimal rate for every task. Read the NeurIPS 2019 paper.
Why learning-rate schedules matter
A schedule changes the learning rate over training rather than keeping it fixed. Warm-up starts with a lower rate and increases it; decay reduces it as training proceeds; restarts raise it again at planned points. A schedule can help balance faster early progress with more controlled later updates, but its effect should be assessed on the task’s validation metric.
Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallBest Value
- Used Book in Good Condition
Google’s speech-recognition study found that schedule choices affected convergence speed and word-error rate in its experiments. Those results show that scheduling can change both training efficiency and task performance, not that one schedule is best for every model. Read Google’s speech-recognition study.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.How to tune the learning rate
- Choose a sensible starting range. Use an order-of-magnitude range suited to the optimizer and model family. There is no universally best numeric learning rate.
- Run a short logarithmic sweep. Try rates spaced by powers of ten rather than only making tiny incremental changes. Track training loss, validation loss, gradient norms, and signs of instability.
- Find a rate that makes prompt, stable progress. Favor a rate that reduces training loss without sustained oscillation or divergence. A sharp initial drop is not enough if training later becomes unstable.
- Tune the schedule with batch size. Compare warm-up, decay, or restarts using validation metrics, not training loss alone.
- Retune after meaningful changes. Recheck the rate if you change the optimizer, batch size, normalization, architecture, or data preprocessing. Each can change effective step sizes or the curvature encountered during training.
What to compare when choosing a rate
For each candidate rate or schedule, record the measures that match your objective. A rate that lowers training loss fastest may not give the best validation result or the lowest compute cost.
Quick Recap
- Initial loss decrease: Does training begin making useful progress?
- Time or updates to target quality: How quickly does the run reach a chosen validation metric?
- Stability: Does loss settle, oscillate persistently, or diverge?
- Validation performance: Does the model improve on held-out data, not only on the training set?
- Batch-size sensitivity: Does the choice still work after batch size changes?
- Compute cost: How much training time or resource use is required to reach the target?
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

