Free tools Windows power users keep installed
One-click scans. No signup required.
Yes, in a limited and measurable sense. Recent research systems have used AI to modify the code of an agent or its tools, then tested those changes on selected tasks and kept variants that performed better. That can improve an agent’s measured performance; it does not show that the AI has changed its underlying foundation-model weights, become generally more intelligent, or learned to train a smarter foundation model on its own.
What does it mean for an AI to rewrite its own code?
In the best-known recent examples, “self-improvement” refers to an outer process that changes the software around a pretrained model: the agent’s code, tools, or procedures for solving tasks. A foundation model generates or helps make a proposed change; the system then evaluates the modified agent. The model weights themselves can remain fixed throughout.
This distinction matters. Changing an agent’s code may help it use tools, manage context, or organize work more effectively, but it is not the same as changing the model that produces its responses. Nor does a better score on a chosen benchmark establish a general increase in intelligence.
How the Darwin Gödel Machine tests code changes
The Darwin Gödel Machine (DGM), described by Zhang and co-authors in a 2025 paper, is a self-modifying coding-agent system. It selects an agent from an archive, uses a foundation model to produce a modified version, and evaluates that version on coding benchmarks. Versions that compile and retain the ability to edit a codebase can continue through the process.
#1 Best Overall
The archive lets the system explore more than a single chain of changes: it can select earlier agents as starting points and preserve variants for later use. Candidate changes are not assumed to be improvements in advance. Their value is judged by evaluation on the selected tasks.
What the paper measured
The DGM paper reports that its system’s score rose from 20.0% to 50.0% on SWE-bench and from 14.2% to 30.7% on Polyglot. These are experimental coding-benchmark results reported by the paper’s authors, not measurements of general intelligence or evidence that any AI will improve by the same amount.
Rank #2
The authors describe changes including better code-editing tools, long-context management, and peer-review mechanisms. The result is evidence that an agent’s software design can be iteratively adjusted to improve performance on chosen coding evaluations. It is not proof that every rewrite will help, or that the system has found a universally better way to reason.
What DGM does not do
The DGM experiments use frozen, pretrained foundation models and focus on the design of coding agents. The authors explicitly say they did not rewrite training scripts and train a new foundation model: “However, we do not show that in this paper, as training FMs is computationally intensive and would introduce substantial additional complexity, which we leave as future work.” That is a proposed direction, not a demonstrated result.
The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →How newer approaches extend the idea
Later research broadens what can be edited or what kinds of tasks are evaluated. These studies remain bounded experiments: each measures performance under particular procedures and evaluations rather than establishing a general rule that AI systems will continuously make themselves more intelligent.
| System | What it can change | What is evaluated | Reported result and boundary |
|---|---|---|---|
| Darwin Gödel Machine (DGM) | A coding agent’s software, including its tools and other agent-design choices; the foundation model remains frozen in the described experiments. | Coding benchmarks, including SWE-bench and Polyglot. | The 2025 paper reports score increases from 20.0% to 50.0% on SWE-bench and from 14.2% to 30.7% on Polyglot. Training a new foundation model is not demonstrated. |
| DGM-H / HyperAgents | A task agent and the meta agent that modifies it are represented in an editable program, so the improvement procedure itself can evolve. | Experiments reported in coding, paper review, robotics reward design, and Olympiad-level math-solution grading. | Meta reports these experimental domains and says the experiments used sandboxing and human oversight. The results do not establish unconstrained self-improvement or a general safety guarantee. |
| AIDE² | A research agent’s harness: the outer-loop software that shapes how its inner-loop task-solving process works. | AI research and transfer to four held-out benchmarks. | A September 22, 2026 preprint reports seven accepted successive improvements during an autonomous eight-day run and says the resulting agent matched or exceeded a human-engineered agent on the held-out benchmarks. The authors note noise and the cost of additional runs; this is a preprint result, not an established general law. |
HyperAgents: modifying the improvement procedure
Meta’s HyperAgents work, also described as DGM-H, goes beyond changing only a task-solving agent. It places a task agent and a meta agent—the part that proposes changes—in one editable program. In principle, this allows the procedure used to make improvements to be modified as well. Meta reports experiments across coding and several non-coding domains, including paper review, robotics reward design, and grading Olympiad-level math solutions.
Meta states: “All experiments were conducted with safety precautions (e.g., sandboxing, human oversight).” Those are precautions reported for these experiments, not a guarantee that every self-modifying system is safe.
AIDE²: making an AI research agent more efficient
AIDE², described in a preprint posted September 22, 2026, targets the efficiency of an AI research agent. Its outer loop rewrites the agent’s harness to improve the inner loop that solves research tasks. The authors report seven accepted successive improvements in an autonomous eight-day run and transfer to four held-out benchmarks, where they say the agent matched or exceeded a human-engineered one.
Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallBest Value
The authors also identify noise and the cost of additional runs as limitations. Because this is a recent preprint report, its findings should be read as results from that study, not as independently established evidence that autonomous research agents will reliably improve in other settings.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.How to interpret “recursive self-improvement”
In these experiments, “recursive” describes a loop: an agent is modified, evaluated, and potentially used as the next version in further modification. The word describes the process; by itself, it does not mean improvement is uncontrolled, exponential, or guaranteed.
- The target may be the agent, not the model. Editing source code, tools, or a harness does not necessarily change the pretrained model’s weights.
- Improvement is defined by an evaluation. A higher score means better performance on the selected tasks and metric, within the tested setup.
- Evaluation choices shape the result. Benchmark design, task selection, the number of attempts, and the evaluation budget all affect what variants are retained. The DGM paper explicitly treats coding benchmarks as a proxy for coding and self-modification ability.
- Repeated improvement is not proof of general intelligence gains. Performance on coding or other selected tasks does not establish a broad increase in reasoning ability across unrelated settings.
- Reported safeguards have a limited scope. Sandboxing and human oversight were reported in the cited HyperAgents experiments; that does not provide a blanket safety guarantee for other systems.
Does this mean AI will inevitably build a smarter version of itself?
No. The studies show narrower forms of software-level iteration, and the DGM paper does not demonstrate training a new foundation model. In its article “When AI builds itself,” Anthropic cautions: “We are not there yet, and recursive self-improvement is not inevitable.” Anthropic discusses both potential benefits and the possibility that humans could lose control as implications of full recursive self-improvement; those are potential outcomes, not findings established by the benchmark results above.
Earlier work on self-programming AI
A 2022 paper, “Self-Programming Artificial Intelligence Using Code-Generating Language Models,” explored a code-generating model modifying its own source code and properties such as architecture, computational capacity, and learning dynamics. It is useful historical context for the idea. The more recent agent-loop studies provide the relevant measured examples here, and their findings remain specific to the systems, tasks, and evaluation procedures they report.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

