Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Clear out junk files and repair common Windows errors3Scan for outdated or missing drivers - takes under a minuteThere is no universally best learning rate for either SGD or Adam. Start from a documented baseline where one is available, compare clearly different candidate rates on your own task, and choose using validation results and training stability. Adam adapts updates for individual parameters, but it still needs a global learning-rate setting.
What the learning rate controls
The learning rate sets the scale of an optimizer’s parameter updates. A rate that is too large for a particular model and task can make training unstable or prevent useful progress; a rate that is too small can make progress slow within your training budget. The practical choice is the rate that trains reliably and performs well on validation data under a controlled comparison.
SGD uses stochastic gradients with its configured rate and may also use momentum. Adam estimates first and second moments of gradients to adapt updates by parameter, but it also exposes a global learning-rate setting. Their update mechanisms differ, so the same numeric rate should not be treated as interchangeable between the optimizers. Neither method is a universal winner; compare them on the task and metric that matter to you. Kingma and Ba’s Adam paper describes the algorithm and settings used for its tested problems.
Choose a baseline, not a supposed optimum
Adam
An Adam learning rate of 0.001 (1e-3) is a reasonable baseline: it is the default documented in current PyTorch Adam documentation, and it matches the α value Kingma and Ba report as a good default for the machine-learning problems they tested. The paper also reports β1=0.9, β2=0.999, and ε=10-8 for those problems. These are reference settings, not a guarantee of best performance on another model, dataset, or implementation.
Recommended Free Tools
#1 Best Overall
SGD
Choose an explicit initial rate for SGD and test it. The references here do not establish a general-purpose numerical SGD rate, so there is no evidence-based single value to recommend for every task. If you use momentum, record its value along with the learning rate.
Run a controlled comparison
- Define the evaluation. Set aside validation data and choose a metric relevant to the task. Decide what constitutes useful progress and what signals unstable training.
- Hold the experiment steady. Keep the model, initialization, data processing, split, batch size, schedule, and training budget the same across candidate runs. Set seeds where practical; repeat close comparisons if run-to-run randomness makes the result unclear.
- Select candidate rates. Compare a modest set of rates that differ clearly in scale. There is no universally established candidate grid or multiplier: choose values suitable for your setup and record them.
- Watch training and validation. Reject settings that make training unstable or fail to make useful progress. Judge generalization with the validation metric rather than training loss alone, and compare validation outcomes within the same budget.
- Retest the complete configuration. If you change the optimizer, rate, or schedule, evaluate the resulting configuration under the same protocol. Report what you tested rather than calling it optimal beyond that setup.
Decide whether a schedule is useful
A fixed learning rate is not the only option. A schedule can change the rate by epoch or batch; TensorFlow’s guide gives examples including exponential, piecewise-constant, polynomial, and inverse-time schedules. Its documentation says a common pattern is to reduce learning as training progresses, but that does not make any particular schedule best for every task.
Rank #2
- Use scikit-learn to track an example ML project end to end
- Explore several models, including support vector machines, decision trees, random forests, and ensemble methods
- Exploit unsupervised learning techniques such as dimensionality reduction, clustering, and anomaly detection
- Dive into neural net architectures, including convolutional nets, recurrent nets, generative adversarial networks, autoencoders, diffusion models, and transformers
- Use TensorFlow and Keras to build and train neural nets for computer vision, natural language processing, generative models, and deep reinforcement learning
For a validation-responsive option, TensorFlow documents ReduceLROnPlateau, which changes the current rate when validation loss stops improving. In Keras, schedule objects can be passed as an optimizer’s learning-rate argument. Choose between a fixed value, a predetermined schedule, or a validation-responsive change based on the behavior you observe, and compare the final scheduled setup against your baseline. TensorFlow’s training and evaluation guide and Keras learning-rate schedule documentation describe these options.
Compare SGD and Adam on the same terms
When choosing between the optimizers, evaluate each using the same validation split, task metric, and comparable training budget. Consider:
Rank #3
- Validation performance on the metric you care about.
- Stability, including erratic loss or divergence.
- Useful progress within the available steps or compute budget.
- Sensitivity to the starting rate and any schedule.
- Whether the optimizer settings and resource limits make the comparison fair.
A result from one model, dataset, or budget does not establish a general winner. Keep the conclusion tied to the tested configuration.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Record the settings so the result is reproducible
Learning-rate defaults and APIs can change by framework version. For each run, report the framework and version, optimizer, initial rate, any optimizer parameters such as momentum or Adam betas, schedule and its parameters, batch size, training budget, and validation criterion. When a rate is changed in response to validation behavior, record that rule too.
Quick Recap
Best Value
Rank #4
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

