Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Backpropagation through time (BPTT) trains an LSTM by unrolling its recurrent equations across a sequence, then applying the chain rule backward through each time step. At every step, the gradient splits across the LSTM’s gates and its additive cell-state update. The forget gate can preserve a gradient along the cell-state path when its values remain near one, but LSTMs do not guarantee that gradients will never vanish or explode.

What backpropagation through time does in an LSTM

An LSTM processes a sequence one position at a time. At position t, it uses the current input xt, previous hidden state ht−1, and previous cell state ct−1 to compute new gate values and states. BPTT makes this recurrent process differentiable across time: it treats the sequence of computations as an unrolled graph and propagates the loss gradient backward through that graph.

A common modern formulation is:

ft = σ(Wfxt + Ufht−1 + bf)
it = σ(Wixt + Uiht−1 + bi)
gt = tanh(Wgxt + Ught−1 + bg)
ct = ft ⊙ ct−1 + it ⊙ gt
ot = σ(Woxt + Uoht−1 + bo)
ht = ot ⊙ tanh(ct)

Here, σ is the sigmoid function, ⊙ is elementwise multiplication, and f, i, g, and o are the forget, input, candidate, and output values. The cell state c carries information along the sequence; the hidden state h is the exposed output used by later computation. Frameworks may combine gate calculations into one matrix operation, but the gradient paths are equivalent.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Why the same weights receive gradients from many time steps

The matrices and biases are shared across time: the same W, U, and b parameters are used at each position. During the backward pass, each position contributes a gradient for those shared parameters. BPTT adds those contributions together. For example, a loss at a later position can affect a shared recurrent weight through multiple earlier hidden states, not just through the final operation that used that weight.

How the gradient moves through the gates and cell state

Reverse-mode differentiation starts at positions where a loss is computed. At each step, gradient flows backward from the hidden state through the output gate and the cell-state activation. It then splits at the cell update: one branch goes through the retained previous state, ft ⊙ ct−1; the other goes through the new content, it ⊙ gt.

The cell-state gradient has a direct retention path

Ignoring other branches for a moment, the derivative of the cell update with respect to the previous cell state is ∂ct/∂ct−1 = ft, element by element. Thus, a gradient traveling along this direct path is multiplied by the forget-gate values at each step. If those values stay near one, this path can carry gradient over many steps; if they are smaller, the path attenuates. Other gradient paths through gates and hidden states still contribute, so this direct-path description is not the entire derivative of the network.

The output gate controls how much of the activated cell state appears in ht. The input gate controls how much candidate content is written into ct. The forget gate controls retention of the previous cell state. These gates are differentiable: the backward pass propagates through their sigmoid or tanh functions as well as through their inputs.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Local derivative sketch

Let δct mean the total gradient arriving at the cell state after combining gradient from the next cell state with gradient through the current hidden output. The local gate-preactivation gradients are then:

  • δaf = (δct ⊙ ct−1) ⊙ ft ⊙ (1 − ft)
  • δai = (δct ⊙ gt) ⊙ it ⊙ (1 − it)
  • δag = (δct ⊙ it) ⊙ (1 − gt2)
  • δao = (δht ⊙ tanh(ct)) ⊙ ot ⊙ (1 − ot)

In these expressions, a denotes a gate’s preactivation, and δht is the total gradient arriving at the hidden state. The cell gradient also receives the path through the hidden output: δht ⊙ ot ⊙ (1 − tanh2(ct)). The gradient passed directly to the previous cell state is δct ⊙ ft. Gradients to the previous hidden state flow through all four gate computations because each gate depends on ht−1.

Once a gate’s preactivation gradient is known, the chain rule gives its parameter gradients. For example, the forget-gate contribution to Wf at step t is the outer product of δaf and xt; its contribution to Uf uses ht−1. These contributions are summed across the unrolled steps, as are the corresponding contributions for the other gates.

Why LSTMs can help with vanishing gradients—and what they do not guarantee

In a conventional recurrent network, a gradient propagated backward across many steps repeatedly passes through recurrent transformations and activation derivatives. Depending on their magnitudes, these repeated multiplications can make the gradient shrink toward zero or grow excessively. The 1997 paper Long Short-Term Memory by Sepp Hochreiter and Jürgen Schmidhuber describes this difficulty and introduces a cell design intended to provide a constant-error route. The authors report learning minimal time lags “in excess of 1000 discrete-time steps” in their experiments; that result is not a general guarantee for every LSTM or task.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The additive cell update gives the modern LSTM a comparatively direct path for state and gradient flow. The forget gate makes retention controllable rather than forcing the cell to overwrite its contents at every step. As the original authors put it, “Multiplicative gate units learn to open and close access to the constant error flow.”

This mechanism improves the opportunity to preserve long-range information, but it does not eliminate vanishing or exploding gradients. A forget value consistently below one attenuates the direct cell path over time, and other paths still pass through nonlinearities and learned transformations. Large gradients can also destabilize training. Gradient-norm monitoring and gradient clipping are common engineering responses to exploding gradients; neither makes the architecture immune to every optimization problem.

Initialization can affect early retention behavior. University of Michigan notes on LSTMs explain that a low initial forget value can repeatedly attenuate the cell path, while a positive forget bias makes initial retention more favorable. This is an initialization consideration, not a promise that the trained gate will remain open or that a particular dependency will be learned.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What truncated BPTT changes

Full BPTT differentiates through the complete unrolled sequence. That can require substantial memory and computation for long sequences, since activations needed for the backward pass must be retained or recomputed. Truncated BPTT limits how far the backward graph extends, using a chosen number of steps as its gradient window.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

In a typical truncated setup, the recurrent state can be carried forward into the next segment, while the computational graph is cut at the segment boundary. The next segment therefore starts from a state influenced by earlier inputs, but its loss does not send a direct gradient through that boundary to those earlier computations. Dependencies older than the window cannot receive a direct learning signal through that particular backward pass.

Choose a window long enough to cover the dependency horizon the task needs, while accounting for the additional memory and compute of longer windows. Truncation makes training more manageable for long streams, but it changes the gradient being computed: it is not identical to differentiating the loss through the entire sequence. Hochreiter and Schmidhuber’s 1997 paper also discusses truncating gradients at architecture-specific points while preserving the intended long-term error route in their design.

LSTM and vanilla RNN: what differs during BPTT

Aspect LSTM Vanilla RNN
Gradient-memory path The additive cell update provides a direct path multiplied by the forget gate; additional paths pass through gates and activations. Gradients pass through repeated recurrent transformations and activation derivatives, which can shrink or grow over many steps.
Information flow Forget, input, and output gates control retention, writing, and exposure of cell state. A recurrent update transforms the hidden state at each step without these separate LSTM gate controls.
Full versus truncated BPTT Full BPTT can be costly for long unrolls; truncation limits direct gradient flow across the chosen window. The same full-versus-truncated trade-off applies: truncation reduces the backward horizon but also limits direct learning signals.
Dependency horizon Can preserve a useful state and gradient path over longer spans, but the learned horizon depends on gates, training, and the task. Long-range learning can be difficult when repeated gradient transformations vanish or explode.

Practical checks when training an LSTM

  • Watch gradient norms: unusually large norms can signal unstable updates; consider gradient clipping when appropriate.
  • Check forget-gate behavior and initialization: early retention depends on the gate values, which are influenced by bias initialization.
  • Set the truncation window deliberately: relate it to the time span over which the task requires a direct learning signal, not just to the sequence’s total length.
  • Separate state carry from gradient carry: carrying hidden and cell values between segments does not by itself preserve gradient flow across a detached boundary.
  • Identify the formulation being discussed: the equations above describe a common modern forget-gated LSTM; the 1997 paper is the foundational historical account, not a claim that every present-day implementation has identical details.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.