The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
A self-improving loop usually fails without crashing. Each cycle proposes a change, something judges it as an improvement, and the loop keeps it. The weak point is the judge. If the judge reads only the agent’s own transcript, the loop can end up confirming its own claims, and the cycle count keeps climbing while the measured result stays flat or falls.
This article does not reconstruct a particular personal project. It uses published 2026 studies to explain this failure pattern, what it looks like in measurements, and which controls catch it.
What a “self-improving loop” actually changes
The phrase covers several different things. Between attempts, a loop can persist changes to the agent’s prompt, its harness (the scaffolding that runs the agent, calls tools, and records results), its memory, or its model weights. Any concrete example should say which of these changes, because the failure modes and the fixes differ. The three studies discussed below each examine a different mechanism, so their numbers should not be merged into one rate.
Quick wins for a faster PC:
Repair Windows errors before they cause bigger problemsFix Now →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Clear out junk files and repair common Windows errorsFree Scan →The stagnation problem, measured
Claimed progress versus measured progress
Hyundoo Park and Byungho Choi’s 2026 arXiv preprint, When Do Agent Loops Mistake Stagnation for Progress?, studies a long-running agent-loop testbed. In that testbed, the agent claimed an improvement in every one of 54 cycles. Yet 56 percent of cycles had a measured delta of zero or below.
#1 Best Overall
Treat that figure as a property of this testbed and its setup, not as a general failure rate for agent loops. What it does show is that a loop’s own sense of progress can diverge sharply from what is measured, and that a loop with no independent measurement cannot tell the difference.
What happened when the agent judged itself
The same paper reports that under a self-verdict gate, the loop eroded the best deployed state it had reached by 19 percent. In plain terms, the loop accepted changes that made the agent worse than a version it had already achieved, and it had no mechanism to restore that version. This is one experimental result, with the setup described in that paper.
Rank #2
Why a stronger judge does not fix it
The authors’ argument is about where the evidence comes from, not how capable the judge is. Their abstract states:
“For open-ended objectives whose success signal lives outside the transcript, scaling up the judge is not enough; out-of-band evaluation with real-world access is a structural requirement.”
The scope matters. The claim concerns open-ended objectives where success is external to the transcript: a booking actually completed, a file actually produced, a page actually changed. If the only thing a judge can read is what the agent wrote about its work, a more capable judge is still reading the agent’s account. Out-of-band evaluation means checking the world, or an independent measurement of it, rather than the agent’s description.
Gate every promotion, not just every proposal
Nakajima’s 2026 arXiv preprint describes Regimes, an auditable self-improvement loop demonstrated on the LongMemEval-S benchmark. Its design separates proposing a change from accepting it. A candidate must clear four gates in order:
- Static checks on the proposed change, before anything runs.
- Sandbox execution, so the candidate runs in isolation rather than in the live agent.
- In-sample evaluation on the same data that informed the proposal.
- Held-out validation on data the proposer never used.
The fourth gate is the one that catches overfitting to the loop’s own examples. Consider an illustrative case, not a measured result: a rule for writing memory entries raises the score on the examples used to tune it, but lowers it on unseen questions. Without the held-out gate, the loop promotes the rule because it looks good where it was built. With the gate, the candidate is rejected and the reason is recorded. These gates are controls that reduce this risk. They are not guarantees of performance on new tasks.
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Repair Windows errors before they cause bigger problems3Scan for outdated or missing drivers - takes under a minuteLearn from failed trajectories, not only successes
Sun and co-authors’ 2026 arXiv preprint studies failure-driven, inference-time self-improvement for computer-use agents on OSWorld. The approach diagnoses failed trajectories and proposes changes applied at inference time, with light human verification of those proposals.
Best Value
Failures carry information that a loop can use. The study’s results apply to that OSWorld setup and should not be read as a general claim about computer-use agents.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.How the three approaches compare
The table below lists what each paper’s summary states. Where a source does not describe an axis, the cell says so. The studies test different things, so the table is a way to see what each design covers, not a ranking.
| Axis | Park and Choi (2026), stagnation testbed | Nakajima (2026), Regimes | Sun et al. (2026), OSWorld computer-use agents |
|---|---|---|---|
| What persists between attempts | Not stated in the paper’s abstract | Not stated in the paper’s summary | Inference-time changes, per the study |
| Source of the success signal | Compares in-transcript self-verdict with out-of-band evaluation | Benchmark evaluation on LongMemEval-S, in-sample and held-out | Outcomes in the OSWorld environment |
| Held-out check before promotion | Not stated in the paper’s abstract | Yes, required as the fourth gate | Not stated in the paper’s summary |
| Runs and promotions can be audited | Not stated in the paper’s abstract | Yes, described as auditable | Not stated in the paper’s summary |
| Failed trajectories analyzed | Not stated in the paper’s abstract | Not stated in the paper’s summary | Yes, failures are diagnosed to drive changes |
| Human review of changes | Not stated in the paper’s abstract | Not stated in the paper’s summary | Light human verification of proposed changes |
Because the setups differ, avoid concluding that one design is better from these results alone.
Free tools Windows power users keep installed
One-click scans. No signup required.
An audit checklist for your own loop
Separating proposing from accepting is the practical lesson. To check whether a loop does this, go through the following:
- Name the persisting artifact: prompt, harness, memory, or model weights. A claim of improvement is meaningless until you know what changed.
- Identify the success signal and confirm it does not come only from the agent’s own text.
- Keep the proposer and the grader separate, so that the same run does not both suggest a change and approve it.
- Run candidates in a sandbox before any promotion.
- Hold out data the proposer never sees, and require a candidate to pass it.
- Count cycles with a measured delta of zero or below, not only cycles where the agent reported progress.
- Keep a best-known-good state and re-check it after promotions, so that a series of accepted changes cannot quietly erode it.
- Log every promotion and rejection with the candidate, the metric, and the gate that decided.
- Review analyzed failures and have a person check proposed changes before they become standing behavior.
The Bottom Line
A loop that accepts changes on its own say-so will report progress whether or not progress happened. The fix is structural: measure success outside the agent’s transcript, put candidates through held-out validation before promotion, and keep a record of every decision. Those controls are not perfect, but they are what separates a loop that improves from one that only appears to.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

