Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

An LSTM (long short-term memory) is a recurrent neural-network layer that processes a sequence one element at a time while carrying information forward in two states: a cell state and a hidden state. Three learned gates regulate what information is retained, added, and exposed at each step. This design was introduced to make learning dependencies over longer time lags more tractable, but it does not guarantee that a model will learn any arbitrary long-range relationship.

How an LSTM processes a sequence

At time step t, an LSTM receives the current input vector xₜ, the previous hidden state hₜ₋₁, and the previous cell state cₜ₋₁. It uses the input and prior hidden state to calculate gate values and a candidate update. It then updates the cell state and produces a hidden state for this step. That process repeats for each element in the sequence.

The hidden state carries information from earlier points in the sequence and serves as the step’s output to the next computation. The cell state is a separate memory representation that is updated through the sequence. In a notebook analogy, the cell state is the running page, the forget gate scales what remains, the input gate scales a proposed addition, and the output gate scales what is shown. This is only an analogy for learned vector operations: gates are not literal switches.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What the three gates do

In the standard PyTorch formulation, each gate is a learned function of the current input and previous hidden state. Sigmoid, written σ, produces values between 0 and 1; these values scale information element by element rather than making all-or-nothing decisions.

  • Forget gate (fₜ): scales the prior cell state, controlling how much of each component is retained.
  • Input gate (iₜ): scales the candidate cell content, controlling how much proposed information is added.
  • Output gate (oₜ): scales the cell-derived information exposed as the hidden state.

The candidate cell content is gₜ, calculated with tanh. The update equations show how the parts fit together:

  • iₜ = σ(Wᵢᵢxₜ + bᵢᵢ + Wₕᵢhₜ₋₁ + bₕᵢ)
  • fₜ = σ(Wᵢfxₜ + bᵢf + Wₕfhₜ₋₁ + bₕf)
  • gₜ = tanh(Wᵢgxₜ + bᵢg + Wₕghₜ₋₁ + bₕg)
  • oₜ = σ(Wᵢoxₜ + bᵢo + Wₕohₜ₋₁ + bₕo)
  • cₜ = fₜ ⊙ cₜ₋₁ + iₜ ⊙ gₜ
  • hₜ = oₜ ⊙ tanh(cₜ)

Here, W terms are learned weight matrices, b terms are learned biases, and ⊙ denotes element-wise (Hadamard) multiplication. The cell update combines scaled prior memory with scaled candidate content; the hidden state is a gated view of the updated cell state. See the PyTorch LSTM API for the documented formulation.

Why LSTMs were developed

When recurrent networks are trained by propagating error backward through a sequence, the learning signal can decay across many steps. Sepp Hochreiter and Jürgen Schmidhuber introduced LSTM in 1997 to help preserve error flow over long time lags through a memory mechanism and multiplicative gates. Their paper reported that LSTM could bridge “minimal time lags in excess of 1000 discrete-time steps” in its experimental setting. That is a historical result from the authors’ experiments, not a universal capacity guarantee or a modern benchmark. Read the original paper, “Long Short-Term Memory”.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Where LSTMs are used

LSTMs can be applied when the order of inputs matters and information from earlier steps may help interpret later ones. Documented examples include language modeling and part-of-speech tagging in the PyTorch sequence-models tutorial, and time-series forecasting in a TensorFlow time-series tutorial. These examples illustrate tasks, not evidence that LSTMs outperform other model types for them.

Input shapes and implementation details

In PyTorch, the feature dimension is the final axis. The default layout places sequence length first for batched inputs; setting batch_first=True moves the batch axis to the front. Here, L is sequence length, N is batch size, and H_in is the input feature count.

Input case Documented input shape
Unbatched (L, H_in)
Batched, default (batch_first=False) (L, N, H_in)
Batched with batch_first=True (N, L, H_in)

When initial hidden and cell states are omitted, PyTorch initializes them to zero. The layer also supports multiple recurrent layers, bidirectional processing, and optional projections with proj_size > 0. For bidirectional or projected configurations, check the API’s output and final-state shapes rather than assuming they match a one-direction, unprojected model.

TensorFlow’s tutorial explains that a Keras LSTM cell is wrapped in an RNN layer, which manages state and sequence results. Consult the TensorFlow tutorial and the documentation for the version you use; library APIs and shape details can change.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

How to decide whether an LSTM fits a project

The mechanics alone do not establish whether an LSTM is the best choice for a particular task. There is no basis here for a blanket claim that LSTMs are faster, more accurate, obsolete, or superior to GRUs or Transformers. Compare candidate models on the same task and data, using criteria that reflect the deployment setting:

  • Validation performance on the target task.
  • Sequence length and the dependency structure the model must capture.
  • Training and inference cost under the hardware and latency constraints that matter.
  • How much suitable training data is available.
  • Whether future sequence elements will be available at inference time. A bidirectional LSTM uses context from both directions, so it is unsuitable when future inputs are not yet known.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.