Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

In a Transformer, attention lets each position build a context-aware representation by comparing a query with keys, turning those match scores into weights, and using the weights to combine values. Queries and keys determine which information matters for this calculation; values provide the information that is mixed into the output.

How query, key, and value fit together

Think of attention as a learned lookup over representations—not a literal database search and not human-like focus. A query represents what one position is looking for in the current representation space. Keys provide the features that the query is compared against. Values carry the content that can contribute to the result.

For each query, the mechanism scores its match with available keys. A stronger score leads to a larger normalized weight, so the corresponding value has more influence in the output. The score itself is not the retrieved content: it determines how much of each value is included.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

As Ashish Vaswani and coauthors put it in “Attention Is All You Need” (2017), “An attention function can be described as mapping a query and a set of key-value pairs to an output, where the query, keys, values, and output are all vectors.”

What the scaled dot-product formula does

The original Transformer uses scaled dot-product attention:

Attention(Q, K, V) = softmax(QKᵀ / √dₖ)V

Here, Q, K, and V are matrices of queries, keys, and values; Kᵀ is the transpose of the key matrix; and dₖ is the dimension of the keys. For one query, the calculation proceeds as follows:

  1. Compare query with keys. Dot products in QKᵀ produce a score for each key. A larger dot product indicates a stronger match in the learned representation space.
  2. Scale the scores. Divide by √dₖ. The Transformer paper explains that dot products can grow large as the key dimension increases, pushing softmax toward very peaked distributions and producing small gradients. Scaling moderates the scores before normalization.
  3. Turn scores into weights. Softmax converts the scaled scores into normalized weights. For a given query, the weights sum to one.
  4. Combine values. Multiply the weights by the value vectors and sum them. The result is the attention output for that query: a weighted combination of the available values.

In compact terms, query-key comparisons decide the mixture; the values supply what is mixed. The original paper discusses additive attention, which computes scores with a learned feed-forward scoring function, as another approach. Its Transformer uses scaled dot-product attention instead. The paper’s method and explanation provide the original comparison.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Self-attention and cross-attention use different inputs

Mechanism Where queries come from Where keys and values come from What it lets the model do
Self-attention The sequence representation being processed The same sequence representation Let positions incorporate information from other positions in that sequence
Cross-attention One representation A separate representation Let positions in one representation draw information from another

In self-attention, the query, key, and value vectors are learned projections of the same sequence representation. In the original Transformer, encoder layers use self-attention, while decoder encoder-attention (often called cross-attention) uses decoder representations to query encoder outputs. The 2017 paper describes both arrangements.

Why use multiple attention heads?

A single head performs attention in one learned projection space. Multi-head attention applies several learned projections and attention operations in parallel, then combines their outputs. Vaswani and coauthors’ stated motivation is to let the model attend to information from different representation subspaces at different positions. Heads are therefore parallel learned views of the representations, not a set of manually assigned roles. For the paper’s formulation, see the original Transformer description.

What attention does not tell you on its own

Attention connections alone do not encode sequence order. The original Transformer adds positional encodings to token embeddings so the model can use position information alongside attention. Without that distinction, it is easy to mistake the pattern of connections for an order-aware representation.

Attention weights show how a particular computation distributes weight across values, but they are not, by themselves, a complete or faithful account of everything a model has learned or why it produced an answer. They describe one part of the computation, not a full explanation of the model’s reasoning.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What the original result means—and does not mean

The 2017 paper reported 28.4 BLEU for the Transformer on the WMT 2014 English-to-German translation benchmark, more than 2 BLEU above the existing best results it cited, including ensembles. This is a result reported for that specific historical benchmark in the paper, not a current leaderboard figure. The result is documented by Google Research.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.