Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Transformer architecture is a neural-network design that uses attention to process relationships between tokens instead of relying on recurrent steps. Introduced in 2017 for machine translation, the original Transformer paired an encoder with a decoder. Later models adapted the design: BERT uses an encoder to learn from context on both sides of a token, while generative decoder models usually predict the next token from the preceding sequence.

How Transformer models developed

Before Transformers: sequence models and recurrence

Earlier sequence-to-sequence systems commonly used recurrent neural networks (RNNs), sometimes augmented with attention. An RNN processes a sequence step by step: each step updates a hidden state using the current input and information carried forward from earlier steps. That sequential dependency can limit how much of a training example a system processes in parallel.

In 2017, Ashish Vaswani and seven coauthors proposed a different approach: build sequence modeling around attention and dispense with recurrence and convolution in the core architecture. Their paper describes the Transformer as “based solely on attention mechanisms.” The change made it possible to process the positions in a training sequence in parallel within a layer, rather than waiting for one recurrent step to finish before computing the next. Google Research, “Attention Is All You Need” (2017)

2017: an encoder-decoder for translation

The original Transformer was designed for sequence transduction: take an input sequence, such as a sentence in one language, and generate a related output sequence, such as its translation. Its encoder builds contextual representations of the input. Its decoder generates output tokens one at a time, using both the input representations and the output prefix generated so far.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The paper reported 41.0 BLEU on the WMT 2014 English-to-French translation benchmark after 3.5 days of training on eight GPUs. Google’s account also says the model outperformed recurrent and convolutional models on the reported WMT 2014 English-to-German and English-to-French benchmarks. These are results from the paper’s specific translation experiments, not current rankings across language models or tasks. Google Research, “Attention Is All You Need” (2017); Google Research, “Transformer: A Novel Neural Network Architecture for Language Understanding” (2017)

2018: BERT and the encoder-pretraining branch

BERT demonstrated a different use for Transformer architecture. Rather than generating a translation, it pretrained a Transformer encoder on unlabeled text to produce representations informed by context on both the left and right. A model could then be adapted to tasks such as classification or question answering with a task-specific output layer.

The BERT paper reported new state-of-the-art results on eleven NLP tasks at publication, including GLUE 80.5, MultiNLI accuracy 86.7%, SQuAD v1.1 test F1 93.2, and SQuAD v2.0 test F1 83.1. Those figures describe the paper’s reported evaluations; they should not be read as present-day records. Devlin et al., “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding” (2018)

How self-attention works

Self-attention lets each token representation draw information from other positions in the same sequence. For example, when interpreting a word whose meaning depends on a distant noun, the model can assign attention to that noun rather than relying only on a chain of adjacent recurrent updates.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  1. Represent tokens and positions. The model maps tokens to vectors and adds position information. Attention alone does not inherently encode word order, so the position signal lets the network distinguish different arrangements of the same tokens.
  2. Form queries, keys, and values. For each token, learned projections create a query, a key, and a value. Intuitively, the query represents what the token is looking for, the key represents what another token can offer, and the value carries the information that may be passed along.
  3. Score and mix information. The model compares a token’s query with other tokens’ keys. Higher similarity produces more weight; the weights are normalized and used to combine the corresponding values. The result is a context-aware representation for that position.
  4. Use multiple heads. Multi-head attention performs several such operations with separate learned projections. Their outputs are combined, allowing the layer to represent different kinds of relationships in parallel.
  5. Transform and stabilize each layer. A position-wise feed-forward network further transforms each token representation. Residual connections provide skip paths, while normalization helps stabilize the stacked layers.

In the original encoder-decoder design, the encoder’s self-attention can use information from both sides of an input position. The decoder’s self-attention is masked so that a position cannot see future output tokens. The decoder also has cross-attention, which lets it use the encoder’s representations while generating the output.

What is the difference between the original Transformer, BERT, and generative decoders?

“Transformer” names an architectural family, not one fixed arrangement. The original encoder-decoder, an encoder-only model such as BERT, and decoder-only generative models reuse attention-centered building blocks but differ in what they can attend to, what they are trained to do, and how they are typically used.

Model family Structure and attention Training objective Typical fit Compute and context considerations
Original Transformer Encoder-decoder. Encoder self-attention uses both sides of the input; decoder self-attention is causal and also attends to encoder states. Sequence-to-sequence translation in the original work. Translation and other conditional generation tasks. The 2017 translation experiment reported 3.5 days on eight GPUs for the English-to-French result cited above. No numerical context-window limit is stated in the cited material.
BERT-style encoder Encoder-only, with bidirectional context in its representations. Pretraining to learn bidirectional language representations from unlabeled text, followed by adaptation to tasks. Understanding-oriented tasks such as classification and extraction rather than native free-form generation. No directly comparable compute or numerical context-window figure is stated in the cited BERT paper details.
Generative decoder family Decoder-only, typically using causal left-to-right attention. Autoregressive next-token prediction. Open-ended generation and prompting. Generation proceeds token by token, and long-context computation is a trade-off; no comparable compute figure or context limit is stated here.

The table describes broad design patterns, not every implementation. Later language models do not all use the original encoder-decoder layout, and terms such as “BERT-style” and “decoder family” cover variations in objectives and implementation.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Why did Transformers gain ground over RNNs?

  • More parallel work during training: attention layers can compute representations for multiple sequence positions in parallel. RNNs have a step-to-step dependency that restricts this form of parallelism.
  • Direct paths between positions: self-attention can connect a token to distant positions in a layer, instead of passing information through a long chain of recurrent updates.
  • A flexible basis for different tasks: the encoder-decoder serves input-to-output tasks, an encoder can learn contextual representations, and a causal decoder can generate continuations.

These advantages do not mean attention is free or that every Transformer is faster in every setting. Attention must relate positions to one another, and decoder-based generation still produces tokens sequentially. The architecture changed the balance of computation and made large-scale parallel training more practical; it did not remove all computational costs or make recurrent models impossible to use.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.