Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

Transformers have not been displaced, and “post-transformer” does not mean language models are going away. It describes research into alternatives to the standard attention-centered Transformer stack: selective state-space models such as Mamba, recurrent-style models such as RWKV, long-convolution models such as Hyena, and hybrids that bring attention back into the design. Their potential benefits—such as faster inference or more efficient handling of long sequences—are results reported for particular papers, models, tasks, and implementations, not guarantees that apply everywhere.

What does “post-transformer” mean?

In a standard Transformer, attention lets tokens use information from other tokens in a sequence. That mechanism has helped make Transformers effective across language and other tasks, but its cost and behavior on long sequences have motivated researchers to explore different ways to process information over time.

Alternatives replace or modify parts of that sequence-processing machinery. Some update a compact internal state as tokens arrive; others use long convolutions and gates to mix information across positions. Hybrid designs combine these mechanisms with attention. The aim is not necessarily to eliminate attention, but to find a useful balance among model quality, training cost, inference speed, memory use, and the ability to retain or retrieve relevant details.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The word “emerges” should be read as a research direction, not as evidence of a settled transition. The studies discussed here show promising results in defined experiments; they do not establish that one architecture has won, that Transformers are obsolete, or how widely these alternatives are used in industry.

How do the main approaches differ?

Approach Sequence mechanism What its cited work evaluates or reports
Mamba Selective state-space updates that depend on input content Language modeling, audio, and genomics; the authors report linear sequence-length scaling and fast inference in their experiments. Mamba paper (2023)
RWKV Parallelizable training with inference formulated in recurrent-neural-network style Models up to 14 billion parameters; the authors report performance on par with similarly sized Transformers in their evaluations. RWKV paper (2023)
Hyena Long convolutions interleaved with data-controlled gating Language modeling and operator-speed comparisons at specified sequence lengths; the authors report subquadratic alternatives to attention. Hyena paper (ICML 2023)
Hybrid Mamba Mamba combined with attention Machine translation at sentence and paragraph level; the study reports that attention improved several tested outcomes. WMT 2024 study
RetNet in REM RetNet used in a token-based reinforcement-learning world model with Parallel Observation Prediction Atari 100K benchmark results in a specific world-model study. ICML 2024 paper

Mamba: selective state-space modeling

A state-space model maintains an internal state that is updated as it processes a sequence. The Mamba authors identify a limitation of input-independent state-space dynamics for discrete language: they may not adapt enough to content that should be retained or ignored. Mamba makes model parameters depend on the input, allowing selective propagation or forgetting, and introduces a hardware-aware recurrent algorithm. The authors describe the architecture as having neither attention nor MLP blocks.

In their 2023 paper, the authors report 5× higher inference throughput and describe linear sequence-length scaling. They also report that their 3-billion-parameter Mamba model outperformed same-size Transformers and matched Transformers twice its size on the paper’s pretraining and downstream evaluations. These are the paper’s results for its experiments—not a promise of the same throughput, scaling advantage, or quality on other hardware, implementations, datasets, or tasks. The paper also reports results across language, audio, and genomics; that breadth does not mean the model is established as a production choice in each field.

RWKV: recurrent-style inference

RWKV aims to combine parallel computation during training with an inference process that can be formulated like a recurrent neural network. In the authors’ formulation, inference has constant computational and memory complexity as the sequence grows. That can be attractive when processing a stream, but it is an architectural claim about RWKV’s formulation, not proof that every variant will be faster or use less total memory in every deployment.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The 2023 paper reports training models up to 14 billion parameters and performance on par with similarly sized Transformers in its evaluations. The result is useful evidence that recurrent-style inference can support capable language models at the tested scale. It does not establish equivalence to current Transformers across all model sizes, tasks, or evaluation methods.

Hyena: long convolutions and gating

Hyena replaces attention with a sequence operator built from long convolutions and data-controlled gates. The convolutions are implicitly parameterized, and the design is intended to offer a subquadratic alternative to attention. As with the other approaches, the practical value depends on the task, sequence length, implementation, and hardware—not just the asymptotic description.

In the 2023 paper, the authors report Transformer-quality language modeling on WikiText103 and The Pile with a 20% reduction in training compute at sequence length 2k. They also report Hyena-operator speedups against highly optimized attention: 2× at sequence length 8k and 100× at 64k. Those figures describe the paper’s specified comparisons. They should not be read as end-to-end model speedups or as results guaranteed for arbitrary workloads.

Do these models prove attention is obsolete?

No. A 2024 machine-translation comparison tested RetNet, Mamba, and hybrid Mamba models on sentence- and paragraph-level datasets. Mamba was highly competitive with Transformers in those experiments, while adding attention improved translation quality, robustness to sequence-length extrapolation, and named-entity recall. The authors summarize the finding this way: “Further analysis show that integrating attention into Mamba improves translation quality, robustness to sequence length extrapolation, and the ability to recall named entities.” The WMT 2024 study is a reminder that a new sequence mechanism and attention need not be an either-or choice.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

That result matters because benchmark quality is not only about processing more tokens cheaply. A system may also need to preserve exact details, generalize to longer sequences than it saw in training, or retrieve a name mentioned earlier. In the tested translation setting, the hybrid design improved several of those outcomes. It does not prove that hybrids are best for every task, but it makes a simple “attention is finished” conclusion untenable.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Where have these ideas been tested beyond language generation?

A 2024 ICML study used RetNet in REM, a token-based reinforcement-learning world-model agent, and added Parallel Observation Prediction. On Atari 100K, the authors report 15.4× faster imagination than prior token-based world models in their study and superhuman performance on 12 of the 26 games. These are results from that benchmark and comparison; they show an application beyond language generation, not widespread deployment of RetNet or a general advantage for reinforcement-learning systems.

Mamba’s paper also reports experiments across audio and genomics alongside language. Together, these studies illustrate that alternative sequence-processing mechanisms are being explored in more than one domain. They do not show that all such fields have adopted them or that one mechanism transfers equally well across modalities.

How should you interpret the efficiency claims?

Terms such as “linear scaling,” “constant memory,” and “100× faster” can sound like universal guarantees. They are not. Scaling describes how some cost changes as sequence length increases under a model or analysis; measured speed depends on the specific operator, implementation, hardware, batch size, and workload. An advantage on one benchmark or sequence length may shrink, disappear, or come with a quality trade-off elsewhere.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Check what was measured. Operator speed, inference throughput, training compute, and end-to-end task performance are different quantities.
  • Keep scale matched. A comparison is most informative when model sizes and evaluation conditions are comparable. The RWKV and Mamba papers report particular scale comparisons; those should not be generalized beyond them.
  • Look at the task and sequence length. Hyena’s reported speedups differ by sequence length, and the translation study evaluates sentence- and paragraph-level inputs.
  • Separate benchmark evidence from adoption. These cited studies provide model- and task-specific results, not a reliable industry-wide statistic for adoption of post-Transformer models.

Which architecture could substitute for a Transformer?

There is no evidence here for a universal substitute. Mamba, RWKV, and Hyena explore distinct ways to reduce dependence on standard attention; the translation results show that retaining some attention can help; and the RetNet world-model study demonstrates a research application outside language generation. Which design makes sense depends on the workload and on whether it delivers the required quality, context behavior, throughput, and memory use in a particular implementation.

The most accurate description is a broader architecture search. Researchers are testing alternatives and combinations, while Transformers remain an important reference point. The cited papers make a case for continued experimentation—not for declaring a post-Transformer era already arrived.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.