Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

Richard Sutton’s “bitter lesson” is that, across several histories of AI, general methods able to benefit from more computation—especially search and learning—have tended to outlast systems built around researchers’ hand-coded understanding of a domain. It is a retrospective argument, not a claim that expertise or engineering is useless. The “26 Words” framing in the title is editorial: Sutton did not label his thesis that way.

What did Sutton mean by the bitter lesson?

In his essay, dated March 13, 2019, Sutton frames the history of AI as a 70-year pattern. His concise statement is: “The biggest lesson that can be read from 70 years of AI research is that general methods that leverage computation are ultimately the most effective, and by a large margin.” Read Sutton’s essay.

The contrast is between building a system around human insight into a particular task and using broadly applicable procedures that can improve with additional computation. Sutton’s claim is not simply that computers get faster or that more compute automatically solves a problem. His point is that approaches such as search and learning can keep gaining from computation as it becomes less costly, while a method tied closely to a researcher’s current understanding may hit limits or make later scaling harder.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Sutton identifies search and learning as the two methods that appear to scale in this way. In this context, search explores possible actions or solutions; learning adjusts a system based on data or experience. Both can use computation to examine or improve many possibilities without requiring a person to encode every useful rule in advance.

How the two approaches differ

Question Domain-specialized approach General method that leverages computation
What is built in? Researchers’ explicit knowledge about a particular domain or task. Procedures such as search or learning that can be applied without encoding the full domain solution.
What drives improvement? Further task-specific insight and design. More computation applied to search, learning, or both.
What is the scaling question? Whether additional specialist rules continue to help rather than constrain progress. Whether additional computation can keep improving results on the task.
How should the evidence be read? As one side of a historical contrast in Sutton’s examples, not proof that every specialist technique fails. As a recurring pattern Sutton argues for, not a universal performance guarantee.

This comparison describes Sutton’s argument; it is not a scorecard that predicts the winner for every AI system. Practical systems can combine general methods with useful domain knowledge and engineering.

What examples does Sutton use?

Chess

Sutton points to the 1997 methods that defeated world champion Garry Kasparov, emphasizing massive, deep search rather than an approach centered on reproducing human understanding of chess. The example illustrates his claim that computation-intensive search can surpass hand-built expertise in a defined task.

Go

In Sutton’s account, a similar shift came later in Go. Search and learning from self-play were central to the change. The example broadens the argument beyond chess: the methods need not encode a human expert’s complete account of how to play.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Speech recognition

Sutton contrasts early systems based on human knowledge of linguistics and articulation with statistical approaches. He then describes deep learning as a later step that used more computation and large training sets. His point is about the direction of the methods he reviews, not a claim that linguistic knowledge has no use in speech technology.

Computer vision

For vision, Sutton describes earlier approaches involving edges, generalized cylinders, and SIFT features, then contrasts them with deep-learning networks using convolution and certain invariances. This example supports his broader pattern without establishing that all earlier techniques disappeared from practice or that every vision problem is best solved the same way.

What the bitter lesson does—and does not—claim

The essay is a retrospective argument drawn from selected histories, not a controlled statistical study or proof covering every area of AI. Sutton’s “70 years” is his framing of the period he reviews, not an independently measured statistic. He argues that methods able to exploit increasing computation have repeatedly gained an advantage in his examples.

  • It does say: Researchers’ domain knowledge can produce short-term progress, but approaches built around that knowledge may plateau or impede later gains when more general, computation-intensive methods become practical.
  • It does not say: Human knowledge, domain expertise, or all hand-engineering is futile.
  • It does not establish: That compute alone explains every AI advance, that computation will always become cheaper, or that a general method will outperform a specialist approach on every task.

The useful question is not whether to use expertise at all. It is whether a design gives a system room to improve through scalable procedures—or relies so heavily on a fixed set of human assumptions that those assumptions become a ceiling.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Where to read more

Sutton’s essay is the direct source for the argument and examples: “The Bitter Lesson”. For readers who want a textbook on reinforcement-learning ideas and algorithms, rather than commentary on the essay, Sutton and Andrew G. Barto’s Reinforcement Learning: An Introduction, second edition, is listed by MIT Press. The publisher page gives a November 13, 2018 publication date and hardcover ISBN 9780262039246.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.