Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For a scalar loss, mark the input or model parameters with requires_grad=True, calculate the loss, then call backward() and read the resulting leaf-tensor .grad. For other needs, choose torch.autograd.grad for a returned gradient or vector-Jacobian product, and torch.func transforms for directional derivatives, full Jacobians, or Hessians.

How PyTorch calculates derivatives

PyTorch’s autograd engine records tensor operations as they execute and applies the chain rule to differentiate that computation. As the official automatic differentiation tutorial puts it, “To compute those gradients, PyTorch has a built-in differentiation engine called torch.autograd.”

Set requires_grad=True on tensors for which you need derivatives. In typical training, parameters are leaves of the computation graph; operations involving them produce the loss, and reverse-mode differentiation propagates the loss derivative back to those leaves.

Calculate a scalar gradient with backward()

Use backward() when the output is a scalar, such as a training loss, and you want gradients accumulated into leaf tensors’ .grad attributes.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import torch

x = torch.tensor(2.0, requires_grad=True)
y = x**3
y.backward()
print(x.grad)  # tensor(12.)

Here, y = x³, so its derivative at x = 2 is 12. In a training loop, clear parameter gradients before each independent backward pass when you want a fresh gradient: .backward() adds to existing .grad values rather than replacing them.

Choose between backward() and autograd.grad

Both use autograd, but they serve different output patterns. backward() writes accumulated derivatives into leaf .grad fields. torch.autograd.grad returns derivatives directly, which is convenient for input gradients or calculations that should not update those fields.

Need Use Result
Accumulate a scalar-loss gradient on leaf tensors loss.backward() Gradient is stored in each relevant leaf’s .grad
Obtain a derivative as a return value torch.autograd.grad(output, inputs) Tuple of gradients; does not accumulate them into the inputs’ .grad fields
x = torch.tensor(2.0, requires_grad=True)
y = x**3
dx, = torch.autograd.grad(y, x)
print(dx)  # tensor(12.)

Autograd normally frees the graph after computing derivatives. Use create_graph=True when you need to differentiate the returned derivative again; retaining the original graph is not a routine requirement.

Differentiate a non-scalar output: VJP or JVP

A vector or tensor output has multiple component derivatives, so a derivative request needs a direction or weighting. torch.autograd.grad accepts this weighting as grad_outputs and computes a vector-Jacobian product (VJP), not the full Jacobian.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
y = f(x)  # y may be a vector or higher-rank tensor
v = torch.ones_like(y)
(jt_v,) = torch.autograd.grad(y, x, grad_outputs=v)

This returns vᵀJ, where J is the Jacobian of f with respect to x. Use a different v to weight output components differently. For a directional product in the other orientation, Jv, use torch.func.jvp.

Compute a full Jacobian or Hessian

Use torch.func.jacrev(f) or torch.func.jacfwd(f) when you genuinely need every entry of the Jacobian. A full matrix can be large, particularly when both input and output tensors have many elements. jacrev supports chunk_size to compute rows in pieces when memory is a constraint.

As a practical starting point, reverse mode is often attractive when a function has fewer outputs than inputs; forward mode can be preferable when outputs outnumber inputs. These are shape-based guidelines, not universal speed guarantees. Benchmark the actual function, device, and shapes, and verify that its operators support the chosen transform.

For second derivatives, torch.func.hessian(f) computes a Hessian. If only a directional second derivative is needed, a Hessian-vector product may avoid materializing the entire matrix.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The torch.func API offers composable transforms including grad, vjp, jvp, jacrev, jacfwd, hessian, and vmap. The official API reference describes it as “JAX-like composable function transforms for PyTorch.” The reference labels the API beta and notes incomplete operator coverage, so check compatibility with the PyTorch version and operations in your program.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Pick the derivative operation for the job

What you need Typical choice What it computes
Scalar loss gradient on parameters backward() Accumulates gradients into leaf .grad values
Gradient returned directly autograd.grad or torch.func.grad Derivative without writing to input .grad fields
Output-weighted derivative autograd.grad with grad_outputs, or torch.func.vjp VJP, vᵀJ
Input-direction derivative torch.func.jvp JVP, Jv
Complete Jacobian torch.func.jacrev or torch.func.jacfwd All output-by-input derivatives
Complete second-derivative matrix torch.func.hessian Hessian of a scalar-valued function

Validate gradients and custom operations

If you implement a custom torch.autograd.Function, define backward() for reverse-mode differentiation. To use torch.func transforms, provide the relevant transform methods: vmap() for vmap and jvp() for forward-mode JVP. Compositions such as jacrev, jacfwd, and hessian may require multiple transform-compatible methods. When possible, compose these methods from PyTorch operations.

Test a custom gradient with torch.autograd.gradcheck; use gradgradcheck when second derivatives matter. Gradcheck compares the analytical Jacobian from autograd with a numerical finite-difference Jacobian. PyTorch’s gradcheck mechanics note explains that “The analytical version uses our backward mode AD while the numerical version uses finite difference.” Numerical tolerances and the tested point matter, so passing checks supports correctness for the tested inputs but does not establish it everywhere. The check also has separate handling for complex-valued inputs.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.