For some heavily studied AI tasks, yes: algorithms have improved the amount of capability researchers can get from a given amount of compute faster than the hardware trend associated with Moore’s Law. OpenAI’s 2020 ImageNet analysis found a 44-fold reduction in the compute needed to reach AlexNet-level performance, compared with an 11-fold improvement attributed to Moore’s Law over the same comparison period. That is strong evidence for a task-specific efficiency gain—not proof that algorithms always improve faster than chips, or that semiconductor progress no longer matters.
What does “outpacing Moore’s Law” mean for AI?
Moore’s Law is commonly used as a rough historical benchmark for hardware progress: transistor density, and often the effective capability of computing hardware, has tended to double on a timescale of about two years. It is not a guarantee that every new chip doubles useful AI performance, nor is it a precise law of current hardware development.
Algorithmic efficiency asks a different question: how much training or inference compute is needed to achieve a defined level of performance? Better architectures and optimizers can reduce the operations required; improved training procedures and higher-quality data can produce better results from the same compute budget. The result is an increase in effective compute: a fixed hardware budget can accomplish more of a particular task.
This is not the same as saying AI capability is driven by algorithms alone. OpenAI’s 2018 account identifies three contributors: algorithmic innovation, data—including supervised data and interactive environments—and compute available for training. The factors interact: faster hardware can make more demanding methods practical, while more efficient methods can extend what existing hardware can do.
#1 Best Overall
What is the strongest evidence that algorithms can beat the hardware trend?
OpenAI’s 2020 analysis examined the compute required to reach AlexNet-level performance on ImageNet. It estimated that the compute needed had fallen 44-fold from the 2012 baseline. For the same comparison period, the analysis said Moore’s Law would imply an 11-fold improvement. In that benchmark and comparison, algorithmic efficiency advanced faster than the hardware-efficiency yardstick.
| Measure | Reported result | What it describes |
|---|---|---|
| Compute reduction to reach AlexNet-level ImageNet performance | 44-fold from the 2012 baseline | OpenAI’s 2020 estimate of the compute needed for that benchmark result |
| Moore’s-Law improvement over the same comparison period | 11-fold | OpenAI’s 2020 hardware-trend comparison; the exact end date is not stated in the summary |
| ImageNet algorithmic-efficiency doubling time | 16 months | OpenAI’s 2020 estimate for the analyzed ImageNet trend |
OpenAI summarized its conclusion cautiously: for AI tasks with substantial research effort or compute investment, algorithmic efficiency might outpace hardware efficiency. It also noted that gains from hardware and algorithms multiply and can be on a similar scale over meaningful periods. The ImageNet result therefore supports a qualified claim about effective compute, not a universal rate for every AI application.
Rank #2
How do compute use and language-model estimates compare?
Efficiency gains do not necessarily mean that the field uses less compute overall. Researchers may spend the savings on larger models, more training, or more ambitious tasks. Historical estimates show how quickly compute demand for notable runs has also risen:
| Measure | Reported result | Source and scope |
|---|---|---|
| Doubling time for compute used in the largest AI training runs | 3.4 months | OpenAI, 2018; runs in its series after about 2012 |
| Compute increase from AlexNet to AlphaGo Zero | More than 300,000 times | OpenAI, 2018; its series of training runs |
| Doubling time for training compute of notable AI models | Roughly five months | Stanford HAI / Epoch AI, 2025 |
| Estimated effective-compute doubling time from language-model algorithmic progress | 5 to 14 months | Epoch AI, 2023; the range reflects estimation choices |
These figures are not interchangeable. Training-run compute growth measures how much compute researchers put into selected runs; algorithmic efficiency estimates how much compute is needed for a given level of performance. The periods, model sets, benchmarks, and methods differ, so the numbers do not form a single contest with one winner. Epoch AI’s 2023 analysis also concluded that pretrained language-model performance exceeded what it expected from simply increasing computing resources, but that observation does not make performance independent of compute or data.
Why does AI efficiency improve—and why can costs still rise?
Efficiency can come from several parts of the development process:
- Architectures and optimizers: A model design or optimization method can reach a target with fewer operations.
- Training recipes and data: Better procedures or more useful data can improve performance without increasing the compute budget.
- Hardware–algorithm interaction: A faster accelerator can make a method practical, while a more efficient method can extract more work from a given accelerator.
- Scaling investment: If researchers reinvest efficiency gains in larger training runs, total compute demand can grow even as the cost of a fixed capability falls.
For scale, Stanford HAI’s 2024 estimates put the compute cost of training GPT-4 at $78 million and Gemini Ultra at $191 million. These are estimates of compute cost, not complete development budgets or a direct comparison of the models’ efficiency. They illustrate why lower compute requirements for a particular target do not guarantee that frontier training becomes cheaper in absolute terms.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.How should you judge claims that AI is “outpacing” Moore’s Law?
Before comparing a claimed algorithmic gain with a hardware trend, check that both refer to a comparable task and period. In particular, ask:
- What capability is held fixed? Efficiency is meaningful only relative to a specified benchmark and performance threshold.
- Is the figure about training or inference? The compute to train a model and the compute to use it answer different questions.
- What is being measured? A reduction in compute for one benchmark is not the same as a gain in general AI capability.
- What assumptions apply? Data quality, hardware, model family, and training procedure can affect the estimate.
- Can the result be reproduced? Estimates based on different model sets or methods may produce different efficiency trends.
The practical takeaway is a shifting cost/performance frontier: for some well-resourced, intensively studied tasks, algorithmic progress has delivered larger effective-compute gains than the Moore’s-Law comparison. The size and pace of that advantage depend on what task, model family, time window, and efficiency measure are being compared. Hardware progress remains part of the story because AI capability is built from algorithms, data, and compute together.
Quick Recap
Best Value
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

