Free tools Windows power users keep installed
One-click scans. No signup required.
iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
Gradient descent is an optimization algorithm that adjusts a machine-learning model’s parameters to reduce a selected objective, usually its training loss. It repeatedly measures how the loss changes as parameters change, then moves them in the direction that should lower the loss. The learning rate controls the size of each move.
What gradient descent means in AI
A model makes predictions using adjustable values called parameters, such as a neural network’s weights. A loss function measures how far those predictions are from the desired results. Gradient descent uses the loss’s gradient—the direction and rate of the steepest local increase—to decide how to change the parameters to minimize that objective.
For parameters θ and objective J(θ), a standard update is:
PC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchθ ← θ − α∇J(θ)
- θ represents the model parameters.
- J(θ) is the selected objective, often a training loss.
- ∇J(θ) is the gradient: how the objective changes with respect to the parameters.
- α is the learning rate, or step size.
The minus sign matters: the gradient points toward local increase, so subtracting it moves in the direction of local decrease. This is a local optimization step, not a guarantee that training will find the globally best possible parameters. Stanford’s CS229 lecture notes describe the cost-minimization framing and update rule.
#1 Best Overall
- Use scikit-learn to track an example ML project end to end
- Explore several models, including support vector machines, decision trees, random forests, and ensemble methods
- Exploit unsupervised learning techniques such as dimensionality reduction, clustering, and anomaly detection
- Dive into neural net architectures, including convolutional nets, recurrent nets, generative adversarial networks, autoencoders, diffusion models, and transformers
- Use TensorFlow and Keras to build and train neural nets for computer vision, natural language processing, generative models, and deep reinforcement learning
How a gradient-descent training step works
- Make predictions. Run examples through the model using its current parameters.
- Calculate the loss. Compare the predictions with the target results using the chosen loss function.
- Calculate the gradient. Determine how changing each parameter would affect the loss.
- Update the parameters. Subtract the learning rate multiplied by the gradient from the current parameters.
- Repeat and monitor. Continue with further examples or batches, watching how the loss changes and deciding whether progress is sufficient.
This separates two jobs that are often mentioned together in neural-network training: backpropagation calculates gradients by applying the chain rule through the network; the optimizer uses those gradients to update parameters. Backpropagation is a way to compute the information gradient descent needs, not another name for the update algorithm. Stanford’s Deep Learning Cheatsheet summarizes the relationship between weight updates, learning rate, and backpropagation.
What the learning rate changes
The learning rate scales each update. If it is too small, training may make slow progress. If it is too large, updates can overshoot lower-loss values or oscillate, making training unstable or preventing it from settling. A loss curve can help show whether progress is continuing or flattening; Google’s Machine Learning Crash Course explanation of gradient descent uses linear regression to illustrate this idea.
Rank #2
There is no universally correct fixed number of updates that guarantees success. Whether training converges depends on the objective’s shape, the update method, and hyperparameters such as the learning rate. A flattening loss curve is one useful signal to inspect, rather than proof by itself that the best possible solution has been found.
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Fix the driver behind crashes, sound loss and screen glitches3Repair Windows errors before they cause bigger problemsBatch, stochastic, and mini-batch gradient descent
These variants differ in how many training examples contribute to one parameter update. Here, “batch gradient descent” means an update computed from the full training set; some materials use “batch” more broadly for any selected group of examples.
| Variant | Examples used per update | Update trade-off |
|---|---|---|
| Batch gradient descent | The full training set | Uses a broader data-based gradient, but each update requires processing more examples. |
| Stochastic gradient descent (SGD) | One example | Each update is less expensive, but the gradient is noisier. |
| Mini-batch gradient descent | A subset of examples | Balances the full-set and single-example approaches; commonly used for neural-network training. |
The right trade-off depends on computational cost, gradient noise, memory, and throughput needs. Stanford CS229’s lecture notes discuss gradient descent and stochastic gradient descent, while its deep-learning cheatsheet covers neural-network updates.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.What gradient descent does—and does not do
Gradient descent changes parameters to minimize the objective it is given. It does not choose the loss function, change the training data, or guarantee a global optimum. Those choices and outcomes depend on the model, objective, data, and training setup. Keeping the distinction clear makes the phrase “the model learns” more precise: its parameters are being adjusted in response to a selected measure of error.
Quick Recap
Best Value
Rank #4
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

