What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

Yes—but only in a specific mathematical sense: the update rule of a continuous-state modern Hopfield network is equivalent to the attention operation used in transformers. That gives attention an interpretation as associative retrieval. It does not mean that every part of a Transformer is a Hopfield network, or that attention provides durable memory across separate inputs.

What is equivalent?

In transformer attention, a query is compared with keys to produce similarity scores. A softmax turns those scores into weights, which are used to combine the corresponding values. In the modern Hopfield formulation studied by Ramsauer and colleagues, an update retrieves or combines stored patterns using a corresponding softmax-weighted operation.

The shared mathematical form is the point: under the paper’s specified formulation, one attention update can be understood as a modern Hopfield update. The authors state this equivalence in Hopfield Networks is All You Need (2020 preprint; published at ICLR 2021). It is an interpretation of an operation, not proof that the systems are identical in every respect.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Which kind of Hopfield network?

The correspondence is with a continuous-state modern Hopfield network. It should not be casually extended to every property of the classical binary Hopfield model. “Hopfield network” can refer to different formulations, and the equivalence depends on the particular states, patterns, and update rule being compared.

Ramsauer and colleagues also report that their modern Hopfield formulation can store exponentially many patterns in the dimension of its associative space under its formal setup. That is a result about that model and its assumptions, not a general guarantee that a transformer has unlimited memory or can retain arbitrary facts indefinitely.

Why this does not make the whole Transformer a Hopfield network

The equivalence concerns the attention update, not the complete Transformer architecture. The original Transformer paper proposes an attention-based architecture without recurrence or convolutions, but a Transformer is more than a single attention operation. Its architecture also includes other components, so the operation-level correspondence does not establish that every component performs Hopfield-style retrieval. See Attention Is All You Need for the original architecture.

Nor does the memory interpretation by itself mean that a model keeps persistent memories between unrelated inputs. It describes how an update can retrieve or combine patterns represented in the operation. Claims about persistent storage require separate evidence about the model and its use.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How to read related theoretical results

A separate result about pure-attention architectures does not overturn the update-rule equivalence. Dong, Cordonnier, and Loukas analyze rank loss that grows doubly exponentially with depth in the pure-attention setting they study. The result concerns theoretical behavior in that setting; it is not a blanket conclusion about every practical Transformer, and it does not show that attention and a modern Hopfield update have different mathematical forms. See Attention is not all you need: pure attention loses rank doubly exponentially with depth.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What the newer extension claims

A 2025 NeurIPS abstract for On the Role of Hidden States of Modern Hopfield Network in Transformer reports a generalized correspondence involving an additional hidden-state variable derived from a modern Hopfield network. This is a later extension claim; the abstract alone does not establish how broadly it applies or provide enough detail to assess its assumptions and derivation.

A practical way to assess the claim

  • Ask what is being equated: one attention update, or the full Transformer architecture?
  • Identify the Hopfield model: is the claim about a continuous-state modern network or a classical binary one?
  • Check the assumptions: which states, patterns, and update formulation are specified?
  • Separate implications: a mathematical reinterpretation is not automatically an empirical result or evidence of persistent memory.
  • Keep limitations in scope: a theorem about pure attention in a particular theoretical setting should not be generalized to every practical architecture.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.