Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Gradient descent algorithms differ mainly in how they estimate the gradient, how they use past gradients, and how they adjust each parameter’s step size. Batch, stochastic, and mini-batch describe how much data informs an update; methods such as momentum, AdaGrad, RMSProp, Adam, and AdamW change how that update is calculated. No single optimizer is best for every model: choose candidates based on your task and compare them with a consistent evaluation and tuning process.

What gradient descent does

During training, an optimizer adjusts model parameters to reduce an objective, often a loss function. Gradient descent uses the objective’s gradient to choose an update direction. The learning rate determines the size of that update: an excessively large rate can make training unstable or prevent it from settling, while an excessively small rate can make progress slow. Learning-rate schedules and parameter initialization also affect training, so an optimizer cannot compensate for every issue in a model, dataset, or training setup.

Batch, stochastic, and mini-batch gradient descent

These terms describe how much training data is used to estimate the gradient for each update. The choice trades off computation per update against how frequently the model can update and how much noise is in the estimate. [Ruder’s overview]

Approach Data used for one update Practical trade-off
Batch gradient descent The full training set Each gradient averages over all training examples, but calculating it can be computationally expensive before an update is made.
Stochastic gradient descent One training example Updates can be frequent, but individual-example estimates are noisy, making the optimization path less smooth.
Mini-batch gradient descent A subset of the training set Balances averaging across examples with more frequent updates; mini-batches are common in practical training.

Terminology can be confusing: in machine-learning practice, “SGD” is often used informally for mini-batch training, even though the strict distinction is one example per update. When reading a configuration or paper, check the batch size rather than assuming the word “SGD” means a single-example update.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How optimizers use gradient history and adapt steps

After choosing the data used to estimate a gradient, an optimizer can modify the update using past gradients or adjust step sizes across parameters. These methods have different mechanics and tuning behavior; their names do not imply a universal performance ranking.

Momentum and Nesterov momentum

Momentum uses gradient history to smooth the update direction. This can help reduce oscillation in the path taken during optimization, but it also adds optimizer state and interacts with the learning rate. Nesterov momentum evaluates the gradient at a look-ahead location rather than only at the current parameters. [Ruder’s overview]

AdaGrad

AdaGrad accumulates squared gradients for each parameter and adapts that parameter’s effective step size. Its coordinate-wise adaptation can be useful when gradients are sparse. However, in some deep-learning settings, accumulating squared gradients over the entire history can make effective learning rates become excessively small later in training. That is a conditional limitation, not a guarantee that AdaGrad will fail. [Deep Learning, Chapter 8]

RMSProp

RMSProp uses an exponentially weighted moving average of squared gradients rather than letting all past squared gradients accumulate indefinitely. Older information therefore has less influence, allowing the optimizer to adapt as gradient scales change. Its behavior still depends on settings such as the decay rate and numerical-stability term. [Deep Learning, Chapter 8]

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #3
Sale
Deep Learning (Adaptive Computation and Machine Learning series)
  • Language Published: English
  • Binding: hardcover
  • It ensures you get the best usage for a longer period

Adam

Adam combines moving averages of gradients and squared gradients and applies bias correction in its standard form. It provides adaptive updates, but its results depend on the task, learning rate, schedule, and other implementation and tuning choices. The original paper introduces Adam for stochastic objectives. [Kingma and Ba, “Adam: A Method for Stochastic Optimization”]

AdamW

AdamW decouples weight decay from the adaptive moment estimates. In PyTorch’s documented implementation, weight decay does not accumulate in the momentum or variance. This distinction matters when interpreting regularization settings; check the documentation for the framework and version you actually use, since implementations and defaults are not guaranteed to match. [PyTorch optimizer documentation]

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

How to choose an optimizer for a training task

Treat an optimizer as a candidate to test, not a universal solution. A useful comparison keeps the model, data, evaluation metric, compute budget, and tuning protocol aligned so that differences are interpretable. Ruder’s overview, the Deep Learning textbook, and PyTorch’s documentation cover these methods, but they do not establish a single optimizer that wins across tasks. [Ruder’s overview] [Deep Learning, Chapter 8] [PyTorch optimizer documentation]

  1. Choose an initial optimizer that is supported and appropriate for your training setup.
  2. Set a learning rate and, where relevant, a schedule; monitor training for instability or unusually slow progress.
  3. Compare a small number of plausible alternatives under the same data, model, evaluation metric, and compute limits.
  4. Record optimizer settings and framework version, then select based on the outcome that matters for your task rather than the optimizer’s reputation.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.