Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

Gradient descent is an optimization algorithm that adjusts a machine-learning model’s parameters to reduce a selected objective, usually its training loss. It repeatedly measures how the loss changes as parameters change, then moves them in the direction that should lower the loss. The learning rate controls the size of each move.

What gradient descent means in AI

A model makes predictions using adjustable values called parameters, such as a neural network’s weights. A loss function measures how far those predictions are from the desired results. Gradient descent uses the loss’s gradient—the direction and rate of the steepest local increase—to decide how to change the parameters to minimize that objective.

For parameters θ and objective J(θ), a standard update is:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

θ ← θ − α∇J(θ)

  • θ represents the model parameters.
  • J(θ) is the selected objective, often a training loss.
  • ∇J(θ) is the gradient: how the objective changes with respect to the parameters.
  • α is the learning rate, or step size.

The minus sign matters: the gradient points toward local increase, so subtracting it moves in the direction of local decrease. This is a local optimization step, not a guarantee that training will find the globally best possible parameters. Stanford’s CS229 lecture notes describe the cost-minimization framing and update rule.

#1 Best Overall
Sale
Hands-On Machine Learning with Scikit-Learn, Keras, and TensorFlow: Concepts, Tools, and Techniques to Build Intelligent Systems
  • Use scikit-learn to track an example ML project end to end
  • Explore several models, including support vector machines, decision trees, random forests, and ensemble methods
  • Exploit unsupervised learning techniques such as dimensionality reduction, clustering, and anomaly detection
  • Dive into neural net architectures, including convolutional nets, recurrent nets, generative adversarial networks, autoencoders, diffusion models, and transformers
  • Use TensorFlow and Keras to build and train neural nets for computer vision, natural language processing, generative models, and deep reinforcement learning

How a gradient-descent training step works

  1. Make predictions. Run examples through the model using its current parameters.
  2. Calculate the loss. Compare the predictions with the target results using the chosen loss function.
  3. Calculate the gradient. Determine how changing each parameter would affect the loss.
  4. Update the parameters. Subtract the learning rate multiplied by the gradient from the current parameters.
  5. Repeat and monitor. Continue with further examples or batches, watching how the loss changes and deciding whether progress is sufficient.

This separates two jobs that are often mentioned together in neural-network training: backpropagation calculates gradients by applying the chain rule through the network; the optimizer uses those gradients to update parameters. Backpropagation is a way to compute the information gradient descent needs, not another name for the update algorithm. Stanford’s Deep Learning Cheatsheet summarizes the relationship between weight updates, learning rate, and backpropagation.

What the learning rate changes

The learning rate scales each update. If it is too small, training may make slow progress. If it is too large, updates can overshoot lower-loss values or oscillate, making training unstable or preventing it from settling. A loss curve can help show whether progress is continuing or flattening; Google’s Machine Learning Crash Course explanation of gradient descent uses linear regression to illustrate this idea.

There is no universally correct fixed number of updates that guarantees success. Whether training converges depends on the objective’s shape, the update method, and hyperparameters such as the learning rate. A flattening loss curve is one useful signal to inspect, rather than proof by itself that the best possible solution has been found.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Batch, stochastic, and mini-batch gradient descent

These variants differ in how many training examples contribute to one parameter update. Here, “batch gradient descent” means an update computed from the full training set; some materials use “batch” more broadly for any selected group of examples.

Variant Examples used per update Update trade-off
Batch gradient descent The full training set Uses a broader data-based gradient, but each update requires processing more examples.
Stochastic gradient descent (SGD) One example Each update is less expensive, but the gradient is noisier.
Mini-batch gradient descent A subset of examples Balances the full-set and single-example approaches; commonly used for neural-network training.

The right trade-off depends on computational cost, gradient noise, memory, and throughput needs. Stanford CS229’s lecture notes discuss gradient descent and stochastic gradient descent, while its deep-learning cheatsheet covers neural-network updates.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What gradient descent does—and does not do

Gradient descent changes parameters to minimize the objective it is given. It does not choose the loss function, change the training data, or guarantee a global optimum. Those choices and outcomes depend on the model, objective, data, and training setup. Keeping the distinction clear makes the phrase “the model learns” more precise: its parameters are being adjusted in response to a selected measure of error.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.