Free tools Windows power users keep installed
One-click scans. No signup required.
iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
Richard Sutton’s “bitter lesson” is that, across several histories of AI, general methods able to benefit from more computation—especially search and learning—have tended to outlast systems built around researchers’ hand-coded understanding of a domain. It is a retrospective argument, not a claim that expertise or engineering is useless. The “26 Words” framing in the title is editorial: Sutton did not label his thesis that way.
What did Sutton mean by the bitter lesson?
In his essay, dated March 13, 2019, Sutton frames the history of AI as a 70-year pattern. His concise statement is: “The biggest lesson that can be read from 70 years of AI research is that general methods that leverage computation are ultimately the most effective, and by a large margin.” Read Sutton’s essay.
The contrast is between building a system around human insight into a particular task and using broadly applicable procedures that can improve with additional computation. Sutton’s claim is not simply that computers get faster or that more compute automatically solves a problem. His point is that approaches such as search and learning can keep gaining from computation as it becomes less costly, while a method tied closely to a researcher’s current understanding may hit limits or make later scaling harder.
Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallSutton identifies search and learning as the two methods that appear to scale in this way. In this context, search explores possible actions or solutions; learning adjusts a system based on data or experience. Both can use computation to examine or improve many possibilities without requiring a person to encode every useful rule in advance.
#1 Best Overall
How the two approaches differ
| Question | Domain-specialized approach | General method that leverages computation |
|---|---|---|
| What is built in? | Researchers’ explicit knowledge about a particular domain or task. | Procedures such as search or learning that can be applied without encoding the full domain solution. |
| What drives improvement? | Further task-specific insight and design. | More computation applied to search, learning, or both. |
| What is the scaling question? | Whether additional specialist rules continue to help rather than constrain progress. | Whether additional computation can keep improving results on the task. |
| How should the evidence be read? | As one side of a historical contrast in Sutton’s examples, not proof that every specialist technique fails. | As a recurring pattern Sutton argues for, not a universal performance guarantee. |
This comparison describes Sutton’s argument; it is not a scorecard that predicts the winner for every AI system. Practical systems can combine general methods with useful domain knowledge and engineering.
What examples does Sutton use?
Chess
Sutton points to the 1997 methods that defeated world champion Garry Kasparov, emphasizing massive, deep search rather than an approach centered on reproducing human understanding of chess. The example illustrates his claim that computation-intensive search can surpass hand-built expertise in a defined task.
Rank #2
Go
In Sutton’s account, a similar shift came later in Go. Search and learning from self-play were central to the change. The example broadens the argument beyond chess: the methods need not encode a human expert’s complete account of how to play.
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Repair Windows errors before they cause bigger problems3Scan for outdated or missing drivers - takes under a minuteSpeech recognition
Sutton contrasts early systems based on human knowledge of linguistics and articulation with statistical approaches. He then describes deep learning as a later step that used more computation and large training sets. His point is about the direction of the methods he reviews, not a claim that linguistic knowledge has no use in speech technology.
Computer vision
For vision, Sutton describes earlier approaches involving edges, generalized cylinders, and SIFT features, then contrasts them with deep-learning networks using convolution and certain invariances. This example supports his broader pattern without establishing that all earlier techniques disappeared from practice or that every vision problem is best solved the same way.
What the bitter lesson does—and does not—claim
The essay is a retrospective argument drawn from selected histories, not a controlled statistical study or proof covering every area of AI. Sutton’s “70 years” is his framing of the period he reviews, not an independently measured statistic. He argues that methods able to exploit increasing computation have repeatedly gained an advantage in his examples.
- It does say: Researchers’ domain knowledge can produce short-term progress, but approaches built around that knowledge may plateau or impede later gains when more general, computation-intensive methods become practical.
- It does not say: Human knowledge, domain expertise, or all hand-engineering is futile.
- It does not establish: That compute alone explains every AI advance, that computation will always become cheaper, or that a general method will outperform a specialist approach on every task.
The useful question is not whether to use expertise at all. It is whether a design gives a system room to improve through scalable procedures—or relies so heavily on a fixed set of human assumptions that those assumptions become a ceiling.
Where to read more
Sutton’s essay is the direct source for the argument and examples: “The Bitter Lesson”. For readers who want a textbook on reinforcement-learning ideas and algorithms, rather than commentary on the essay, Sutton and Andrew G. Barto’s Reinforcement Learning: An Introduction, second edition, is listed by MIT Press. The publisher page gives a November 13, 2018 publication date and hardcover ISBN 9780262039246.
Quick Recap
Best Value
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

