Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Repair Windows errors before they cause bigger problemsFix Now →iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
A 77% average pass rate does not mean an AI agent reliably completes 77% of tasks every time. In one AppWorld experiment, a GPT-4.1-backed agent averaged a successful result in 77% of five attempts per task, but succeeded on all five attempts for only 53% of tasks. That 53% is a benchmark measure of repeatability—not a measured production success rate.
What the 77% and 53% figures actually measure
The figures come from the September 8, 2026 arXiv preprint “Closing the Consistency Gap: Self-Evolving Agents That Learn to Stay on Course”, by Evelyn Duesterwald, Benjamin Elder, Lilian Ngweta, Shashanka Ubaru, and Malgorzata Zimon. The authors evaluated ReAct agents on AppWorld’s 168-task test_normal split. Each task was run five times and scored with the benchmark’s standard grader.
- Mean@5: 77%. Across the five attempts, the GPT-4.1 agent averaged a successful result on 77% of attempts.
- Pass^5: 53%. The agent succeeded on every one of the five attempts for 53% of tasks.
- Consistency gap: 24.4 percentage points. This is the reported difference between Mean@5 and Pass^5 for that setup.
The measures answer different questions. Mean@5 captures average success across attempts; Pass^5 asks whether success held on every attempt for a task. A task that succeeds three times and fails twice contributes successes to the average, but does not count as an all-five success.
PC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchWhy an average can hide unreliable task completion
Aggregating successful attempts can make an agent look more dependable than it is on any particular repeated task. If a workflow sometimes succeeds and sometimes fails under the same benchmark task and grading rule, its successful attempts still raise the mean. The average alone does not tell you how many tasks were consistently completed versus how many had mixed outcomes.
#1 Best Overall
This matters in use because people typically need the result of an individual run, not an average across many retries. The paper’s 53% figure is not the chance that any one run will pass, and it does not imply that the five outcomes are independent. It is the fraction of tasks that passed all five runs in this experiment.
How to read repeated-run metrics
For a report using k attempts per task, the paper distinguishes three metrics:
- Pass@k: a task succeeds at least once in k attempts.
- Mean@k: the average fraction of successful attempts across tasks.
- Pass^k: a task succeeds on every one of its k attempts.
The consistency gap is Mean@k minus Pass^k, expressed in percentage points. The paper also defines normalized consistency as Pass^k divided by Mean@k. That ratio can help distinguish consistency from raw capability: when average success is low, the absolute gap has a lower mathematical ceiling. In the paper’s GPT-OSS-120B results, hard tasks had a mean pass rate of just 9.5%, mechanically limiting the gap; normalized consistency was zero for those hard tasks in this evaluation.
Recommended Free Tools
What the experiment does—and does not—show
The study evaluated GPT-4.1 and GPT-OSS-120B with a ReAct agent on one AppWorld test split, using five runs per task. It is evidence that average and all-runs success can differ substantially in that setup; it is not evidence that production agents generally have a 53% reliability rate. The authors say they informally saw similar patterns with other architectures but did not quantify those settings.
Rank #3
The size of the gap also varied by model and task difficulty. For GPT-4.1, the absolute gap rose from 17.5 percentage points on easy tasks to 30.2 points on hard tasks. That pattern did not hold in the same way for GPT-OSS-120B: its low hard-task mean constrained the absolute gap. “Harder tasks always create a larger consistency gap” is therefore not a safe generalization.
The authors discuss uncertain decisions during agent execution as a possible source of run-to-run flips and describe a black-box analyzer that resamples response variability. The result should not be reduced to temperature settings: the paper discusses flips even with temperature-zero decoding, so setting temperature to zero is not a guarantee of deterministic outcomes.
Rank #4
What the guideline intervention changed
The paper tested a method that analyzes variability at agent decision steps, creates targeted natural-language consistency guidelines, stores them as episodic memory, and retrieves them for similar tasks. The analysis and guideline-generation stages are described as offline steps.
Do these 3 things before closing this tab:
1Scan for outdated or missing drivers - takes under a minute2Clear out junk files and repair common Windows errors3Fix the driver behind crashes, sound loss and screen glitches- For GPT-4.1 on the same tasks, Pass^5 rose from 53.0% to 69.0%, a 16-percentage-point increase. Mean@5 rose by 3.6 points rather than declining.
- On similar-task generalization, GPT-4.1 Pass^5 increased by 13 percentage points.
- For GPT-OSS-120B, the baseline was 34% Mean@5 and 10% Pass^5; the reported same-task Pass^5 gain was 6 percentage points.
These are results from the paper’s AppWorld experiments, not guaranteed improvements for a deployed agent or a result independently reproduced by another team.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.How to test an agent for repeatability
A useful local evaluation can follow the logic of the paper’s metrics. The steps below are a practical inference from those metrics, not a separately validated protocol.
- Freeze the test cases and success rule. Keep the task set and grader consistent across comparisons so that a changed score is not caused by different cases or grading.
- Run each task multiple times. Record every outcome for each task, rather than keeping only a single attempt or an overall pass count.
- Report both the average and the all-runs rate. State the number of repeats and give Mean@k alongside Pass^k; include Pass@k if “succeeds at least once” is relevant to your use case.
- Make comparisons conditional on evaluation details. When comparing results, identify the benchmark and split, model and agent architecture, repeat count, grader and success definition, task-difficulty mix, and whether a result is a baseline, an intervention on the same tasks, or generalization to similar tasks.
Repeated-run evaluation is particularly useful when a task must work reliably without a human choosing the best result from several retries. An average remains informative, but it should not stand in for a measure of consistency.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

