Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

Encoder-only, decoder-only, and encoder-decoder Transformers use the same core attention calculation; they differ in which positions can exchange information and where each position gets its queries, keys, and values. That distinction determines whether a model is suited to representing a complete input, generating a continuation, or producing an output conditioned on a separate input.

What attention calculates

Scaled dot-product attention takes query, key, and value matrices and computes:

Attention(Q, K, V) = softmax(QKᵀ / √dₖ)V

The query-key product, QKᵀ, scores how closely each query matches each key. Dividing by the square root of the key dimension, dₖ, scales those scores before softmax. Softmax converts each row of scores into weights that sum to one; multiplying those weights by V produces a weighted combination of value vectors.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

In self-attention, queries, keys, and values are learned projections of the same sequence representation. A mask can change which query-key connections are allowed: unavailable connections receive a prohibitive score before softmax and therefore zero weight.

Why use multiple heads?

Multi-head attention repeats the operation with multiple learned query, key, and value projections. It concatenates the head outputs and applies a further projection. Heads can learn different relationships among positions, but that does not guarantee that each head has a distinct, easily interpretable linguistic role.

How the three architectures differ

Architecture Attention pattern What a position can use Common task pattern Examples
Encoder-only Bidirectional self-attention Other positions on either side within the input Representing or classifying a complete input BERT-like encoders
Decoder-only Causal self-attention Its own position and earlier positions; later target positions are masked Next-token prediction and autoregressive generation GPT-like causal language models
Encoder-decoder Bidirectional encoder self-attention, causal decoder self-attention, and decoder cross-attention The decoder uses earlier target tokens and can consult encoded source positions Conditional sequence-to-sequence tasks, such as translation The original Transformer; T5 and BART are common examples

These are common patterns, not immutable rules for every implementation. Hugging Face documents that a causal decoder model can be run with bidirectional attention for a particular use; changing that attention mode does not turn its block architecture into an encoder. Hugging Face’s Attention Interface documentation distinguishes the configured attention behavior from the model architecture.

How causal and bidirectional attention differ

Bidirectional attention uses the full input

An encoder-only model processes the supplied sequence so each token’s contextual representation can depend on tokens to its left and right. This is useful when the complete input is available and the task is to understand or represent it—for example, classification. Google’s overview also describes embeddings as an encoder-only use. Google’s Transformers explanation summarizes these common roles.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #2
Sale
Hands-On Machine Learning with Scikit-Learn, Keras, and TensorFlow: Concepts, Tools, and Techniques to Build Intelligent Systems
  • Use scikit-learn to track an example ML project end to end
  • Explore several models, including support vector machines, decision trees, random forests, and ensemble methods
  • Exploit unsupervised learning techniques such as dimensionality reduction, clustering, and anomaly detection
  • Dive into neural net architectures, including convolutional nets, recurrent nets, generative adversarial networks, autoencoders, diffusion models, and transformers
  • Use TensorFlow and Keras to build and train neural nets for computer vision, natural language processing, generative models, and deep reinforcement learning

Causal attention limits each position to a prefix

A causal mask lets a position attend to itself and earlier positions while blocking later target positions. Without this restriction during next-token training, the model could see the token it is supposed to predict. A decoder-only model therefore learns next-token conditional probabilities given a prefix; at generation time, it appends a predicted token and repeats the process.

For a sequence of tokens, this is a left-to-right factorization: each token is predicted conditioned on the tokens before it. The mask is about preventing access to future target tokens, not about making the attention formula itself a different equation. Hugging Face’s encoder-decoder overview describes causal decoder attention in this generation setting.

What cross-attention does

Encoder-decoder models separate source processing from target generation. The encoder reads the source sequence and produces contextualized states. The decoder uses causal self-attention over the target prefix, then cross-attention to consult the encoder output.

In cross-attention, decoder states supply the queries, while encoder states supply the keys and values. Each output position can use its query to assign weights to relevant source positions and combine their values. The resulting output is conditioned both on the encoded source and on previously generated target tokens. This makes the architecture a natural fit when a task maps one sequence to another, such as translation. The Hugging Face explanation details the source-encoding and autoregressive-decoding path.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Which architecture fits which task?

Choose by the structure of the information and output, rather than assuming one family is best for every task.

  • Represent or classify a complete input: an encoder-only pattern can use context from both directions when the full input is available.
  • Continue a prompt or generate a sequence token by token: a decoder-only pattern’s causal mask matches next-token prediction.
  • Generate a target sequence from a separate source: an encoder-decoder pattern gives the decoder an explicit cross-attention path to source representations.

Useful comparison questions are: Can a position use later tokens, or must it obey a prefix? Is the task to represent a complete input, continue one, or map a source to a target? Does context travel in the same causal sequence or through a separate encoded source? The answers clarify the attention pattern without declaring a universal winner.

What attention costs as sequences grow

In Google’s simplified account, self-attention scaling is O(N² · S · D), where N is context length, S is the number of self-attention layers, and D is the number of heads per layer. The key takeaway is the quadratic sequence-length term in that simplified expression, not a promise about real-world latency.

Actual runtime and memory also depend on dimensions, implementation, hardware, batch shape, and optimization, including the attention kernels used. Sequence length, caching, and hardware can affect practical trade-offs, so architecture labels alone do not establish which model will be faster or cheaper. Google’s course presents the simplified scaling expression and its variables.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What the original Transformer result says—and does not say

In the 2017 paper Attention Is All You Need, Vaswani and coauthors reported 28.4 BLEU on WMT 2014 English-to-German. Its arXiv abstract reports 41.8 BLEU on WMT 2014 English-to-French for a single model trained for 3.5 days on eight GPUs. These are historical results for the paper’s Transformer, not a current comparison of modern LLM architecture families. The paper’s arXiv abstract states those figures.

There is a page-level discrepancy for the English-to-French score: Google Research’s paper page displays 41.0, whereas the arXiv abstract displays 41.8. The figures should not be silently combined or treated as identical; Google Research’s publication page shows the 41.0 figure.

Further reading

For a practical treatment that includes attention mechanisms, Transformer anatomy, self-attention, and encoder, decoder, and encoder-decoder models, see Natural Language Processing with Transformers, Revised Edition by Lewis Tunstall, Leandro von Werra, and Thomas Wolf. O’Reilly lists it as a 408-page English-language book; it is an intermediate-to-advanced practical NLP and Transformers text, not a dedicated mathematical monograph. See the publisher’s book listing and contents.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.