Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

In self-attention, each token uses learned queries and keys to decide how much to draw from other tokens’ value vectors, then combines those values as a weighted sum. In standard full self-attention, every token can compare itself with every token, so doubling sequence length quadruples the pairwise work and the associated memory requirements described for that calculation.

What does attention average?

Attention combines value vectors, not the original words or tokens directly. For each token, the model forms a query, a key, and a value through learned projections. The query is compared with keys to produce scores; a softmax turns those scores into normalized weights; and those weights are applied to the values to produce the token’s output. [Vaswani et al., 2017]

It is a weighted sum rather than an ordinary average with fixed coefficients. The projections are learned during training, while the resulting weights depend on the input: a token can assign different importance to other tokens in different contexts. In multi-head attention, several sets of learned projections perform distinct attention computations whose outputs are combined.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Why does attention scale quadratically with sequence length?

With n tokens in standard full self-attention, each of the n queries can score against all n keys. That produces n × n pairwise interactions, so the standard calculation’s time and memory requirements grow in proportion to n2. NVIDIA’s Transformer Engine 2.15.0 documentation says that, for the described attention calculation, doubling sequence length quadruples runtime and memory requirements. [NVIDIA Transformer Engine 2.15.0 documentation]

#1 Best Overall
GIGABYTE GeForce RTX 5070 Ti Gaming OC 16G Graphics Card, 16GB 256-bit GDDR7, PCIe 5.0, WINDFORCE Cooling System, GV-N507TGAMING OC-16GD Video Card
  • Powered by the NVIDIA Blackwell architecture and DLSS 4
  • Powered by GeForce RTX 5070 Ti
  • Integrated with 16GB GDDR7 256bit memory interface
  • PCIe 5.0
  • WINDFORCE cooling system

The square describes this full pairwise attention calculation; it does not mean every Transformer operation, or every attention variant, necessarily has the same scaling. The original Transformer was proposed as a sequence model based on attention without recurrence or convolution. [Vaswani et al., 2017]

How do memory-optimized and linear attention differ?

Methods that reduce memory use are not automatically methods that change full attention’s sequence-length complexity. The distinction is whether an approach keeps the full attention calculation but handles its intermediate results more efficiently, or reformulates the operation itself.

Rank #2
ASUS Dual Radeon RX 9060 XT 16GB GDDR6 Gaming Graphics Card
  • Axial-tech fans now feature a smaller fan hub that facilitates longer blades and a barrier ring that increases downward air pressure
  • 2.5-slot design allows for greater build compatibility while maintaining cooling performance
  • 0dB technology lets you enjoy light gaming in relative silence
  • Dual BIOS switch lets you toggle between Quiet and Performance BIOS profiles
  • Dual ball fan bearings last up to twice as long as sleeve bearing designs
Approach What changes Sequence-length scaling and memory behavior Evidence and qualification
Standard full self-attention Each query scores against every key in the sequence. Pairwise interactions grow as n2; the standard calculation has quadratic time and memory scaling. NVIDIA’s Transformer Engine 2.15.0 documentation describes runtime and memory requirements quadrupling when sequence length doubles. [NVIDIA Transformer Engine 2.15.0 documentation]
Memory-optimized exact attention Retains the attention calculation while changing how intermediates are handled. NVIDIA describes tiling and recomputation to improve memory efficiency and data movement. Its flash algorithm avoids storing the full softmax matrix for backward computation, saving normalization factors instead. These memory techniques do not, by themselves, make full pairwise attention linear in sequence length. Implementation details are from NVIDIA Transformer Engine 2.15.0 documentation. [NVIDIA Transformer Engine 2.15.0 documentation]
Linear attention Reformulates attention using kernel feature maps and matrix associativity. The cited method states linear, or O(N), sequence-length complexity; it is a different attention formulation, not merely a memory-saving implementation of the standard one. Katharopoulos et al. (2020) report experiments for their Linear Transformers, including up to 4000× faster autoregressive prediction of very long sequences under their reported setups. That result is paper-specific, not a general speed guarantee or evidence of equal quality for every task. [Katharopoulos et al., 2020]

Complexity notation describes how a quantity grows with sequence length; it is not a complete speed benchmark. Practical speed and quality depend on the method, task, sequence lengths, hardware, and evaluation setup. The cited linear-attention result should therefore be read as an experimental finding from that paper, not as a direct performance comparison for current systems.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Where did the Transformer’s attention approach begin?

In their 2017 paper, Ashish Vaswani and coauthors wrote: “We propose a new simple network architecture, the Transformer, based solely on attention mechanisms, dispensing with recurrence and convolutions entirely.” [Attention Is All You Need]

Rank #3
Sale
GIGABYTE GeForce RTX 5060 WINDFORCE OC 8G Graphics Card, Cooling System, 8GB 128-bit GDDR7, PCIe 5.0, Manufactured by NVIDIA, DisplayPort & HDMI - Video Output Interface, GV-N5060WF2OC-8GD Video Card
  • Powered by the NVIDIA Blackwell architecture and DLSS 4
  • Powered by GeForce RTX 5060
  • Integrated with 8GB GDDR7 128bit memory interface
  • PCIe 5.0
  • WINDFORCE cooling system

The paper introduced the Transformer for sequence transduction and reported 28.4 BLEU on WMT 2014 English-to-German and 41.0 BLEU on WMT 2014 English-to-French. These are results reported by the original paper on those historical translation benchmarks, not predictions of how current models will perform on other tasks. [Vaswani et al., 2017]

Quick Recap

Bestseller No. 1
GIGABYTE GeForce RTX 5070 Ti Gaming OC 16G Graphics Card, 16GB 256-bit GDDR7, PCIe 5.0, WINDFORCE Cooling System, GV-N507TGAMING OC-16GD Video Card
GIGABYTE GeForce RTX 5070 Ti Gaming OC 16G Graphics Card, 16GB 256-bit GDDR7, PCIe 5.0, WINDFORCE Cooling System, GV-N507TGAMING OC-16GD Video Card
Powered by the NVIDIA Blackwell architecture and DLSS 4; Powered by GeForce RTX 5070 Ti; Integrated with 16GB GDDR7 256bit memory interface
$1,162.49
Bestseller No. 2
ASUS Dual Radeon RX 9060 XT 16GB GDDR6 Gaming Graphics Card
ASUS Dual Radeon RX 9060 XT 16GB GDDR6 Gaming Graphics Card
0dB technology lets you enjoy light gaming in relative silence; Dual BIOS switch lets you toggle between Quiet and Performance BIOS profiles
$529.00
SaleBestseller No. 3
GIGABYTE GeForce RTX 5060 WINDFORCE OC 8G Graphics Card, Cooling System, 8GB 128-bit GDDR7, PCIe 5.0, Manufactured by NVIDIA, DisplayPort & HDMI - Video Output Interface, GV-N5060WF2OC-8GD Video Card
GIGABYTE GeForce RTX 5060 WINDFORCE OC 8G Graphics Card, Cooling System, 8GB 128-bit GDDR7, PCIe 5.0, Manufactured by NVIDIA, DisplayPort & HDMI - Video Output Interface, GV-N5060WF2OC-8GD Video Card
Powered by the NVIDIA Blackwell architecture and DLSS 4; Powered by GeForce RTX 5060; Integrated with 8GB GDDR7 128bit memory interface
$459.99
SaleBestseller No. 4
GIGABYTE Radeon RX 9070 XT Gaming OC 16G Graphics Card, PCIe 5.0, 16GB GDDR6, GV-R9070XTGAMING OC-16GD Video Card
GIGABYTE Radeon RX 9070 XT Gaming OC 16G Graphics Card, PCIe 5.0, 16GB GDDR6, GV-R9070XTGAMING OC-16GD Video Card
Powered by Radeon RX 9070 XT; WINDFORCE Cooling System; Hawk Fan; Server-grade Thermal Conductive Gel
$814.99
SaleBestseller No. 5
ASUS Dual GeForce RTX 5060 Ti 16GB GDDR7 OC Edition Gaming Graphics Card
ASUS Dual GeForce RTX 5060 Ti 16GB GDDR7 OC Edition Gaming Graphics Card
AI Performance: 767 AI TOPS; OC mode: 2632 MHz (OC mode)/ 2602 MHz (Default mode); Powered by the NVIDIA Blackwell architecture and DLSS 4
$786.37
Best Value
Sale
ASUS Dual GeForce RTX 5060 Ti 16GB GDDR7 OC Edition Gaming Graphics Card
  • AI Performance: 767 AI TOPS
  • OC mode: 2632 MHz (OC mode)/ 2602 MHz (Default mode)
  • Powered by the NVIDIA Blackwell architecture and DLSS 4
  • Axial-tech fan design features a smaller fan hub that facilitates longer blades and a barrier ring that increases downward air pressure
  • A 2.5-slot design maximizes compatibility and cooling efficiency for superior performance in small chassis
Rank #4
Sale
GIGABYTE Radeon RX 9070 XT Gaming OC 16G Graphics Card, PCIe 5.0, 16GB GDDR6, GV-R9070XTGAMING OC-16GD Video Card
  • Powered by Radeon RX 9070 XT
  • WINDFORCE Cooling System
  • Hawk Fan
  • Server-grade Thermal Conductive Gel
  • RGB Lighting

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.