Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

There is no universally best choice between stochastic gradient descent (SGD) and Adam. Adam is a sensible first candidate when gradients are noisy or sparse and can make quick early training progress; SGD, often with momentum, deserves a fair comparison when held-out performance is the priority. Choose by comparing validation results under equally fair tuning and training conditions—not by training loss alone.

How SGD and Adam update model parameters

Both optimizers use gradients to update a model’s parameters, but Adam also tracks gradient history to adjust the scale of each parameter’s update.

SGD uses a learning rate to scale gradient steps

Ordinary SGD takes a gradient step scaled by a learning rate. Momentum variants additionally accumulate update direction over time. Ordinary SGD does not use Adam’s adaptive moment-based scaling.

Adam adapts updates using gradient history

Adam maintains exponential moving averages of gradients and squared gradients. It corrects those estimates for initialization bias, then divides the corrected first moment by the square root of the corrected second moment plus epsilon. This gives parameters individual update scales based on their gradient history. The algorithm and its original authors’ recommendations are described in Kingma and Ba’s 2014 Adam paper.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

When should you use Adam instead of SGD?

Adam is a reasonable candidate when gradients are very noisy or sparse, or when the objective changes over training. Kingma and Ba describe the method as appropriate for non-stationary objectives and very noisy or sparse gradients. Those are contexts identified by the paper, not a guarantee that Adam will be best on every model or task.

Adam can also reduce training loss quickly in some comparisons. That is useful for monitoring optimization, but it does not establish that the model will perform better on unseen data or train faster in wall-clock time.

Rank #2
Sale
Hands-On Machine Learning with Scikit-Learn, Keras, and TensorFlow: Concepts, Tools, and Techniques to Build Intelligent Systems
  • Use scikit-learn to track an example ML project end to end
  • Explore several models, including support vector machines, decision trees, random forests, and ensemble methods
  • Exploit unsupervised learning techniques such as dimensionality reduction, clustering, and anomaly detection
  • Dive into neural net architectures, including convolutional nets, recurrent nets, generative adversarial networks, autoencoders, diffusion models, and transformers
  • Use TensorFlow and Keras to build and train neural nets for computer vision, natural language processing, generative models, and deep reinforcement learning

Which optimizer generalizes better?

Generalization means performance on held-out data, not how low the training loss becomes. In experiments across the models and tasks they evaluated, Wilson and colleagues reported that SGD and SGD with momentum outperformed adaptive methods on development/test sets when methods received the same amount of hyperparameter tuning. Their 2017 paper, The Marginal Value of Adaptive Gradient Methods in Machine Learning, is evidence to test SGD—not proof that it always wins.

The result is bounded by that study’s models, tasks, and tuning protocol. It does not establish a universal ranking for every architecture, dataset, or later optimizer variant. Treat training loss and validation performance as distinct signals: a falling training loss alongside stalled or worsening validation results is a reason to question whether continued optimization is helping the metric that matters.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How to compare SGD and Adam fairly

Use the same data splits, architecture, compute budget, and deployment-relevant evaluation metric for each candidate. Tune each optimizer’s learning rate and schedule comparably, record training and validation behavior separately, and repeat runs if variability could change the result. This workflow is a practical way to apply the task-dependent findings in the cited studies; it is not a checklist prescribed verbatim by either paper.

  1. Set the goal. Choose the held-out metric that best reflects deployment, such as validation accuracy or loss, and establish a baseline.
  2. Choose candidates. Train Adam and SGD; include an SGD-with-momentum candidate when appropriate.
  3. Give them comparable opportunities. Use a comparable hyperparameter search and training budget for each. Do not compare one optimizer’s tuned configuration with another’s untuned defaults.
  4. Track both kinds of progress. Record training loss and validation performance through training. Note whether validation performance plateaus while training loss continues to improve.
  5. Select on reliable held-out performance. Compare the best validation result each candidate achieves under the same protocol. Repeat runs when training variability could change which configuration appears best.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

How to treat Adam’s original recommended settings

Kingma and Ba’s 2014 paper lists tested settings of α = 0.001, β₁ = 0.9, β₂ = 0.999, and ε = 10−8 for the machine-learning problems in that paper. These are historical recommendations from those experiments, not a statement of current defaults in PyTorch, TensorFlow, or another framework, and not a substitute for tuning on your task.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.