Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Clear out junk files and repair common Windows errors3Scan for outdated or missing drivers - takes under a minuteiTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
In a reported test of 170 change-planning goals, Debashish Ghosal says his AI planning system repeatedly ran into three structural problems: unverified prerequisites, unsafe task order, and weak rollback plans. His practical takeaway is to put hard safety rules in deterministic checks and send unresolved plans to a human—not to assume that a plausible plan is a safe one.
What the 170-goal test involved
Ghosal describes PlannerCritic as a workflow in which one large language model drafts a structured plan, deterministic gates check it against hard rules, and a second model critiques plans that pass those gates. A bounded revision loop can make further changes or escalate the plan for human review. The reported sweep covered 40 domains, including identity management, multi-agent operations, site reliability engineering, supply-chain policy, and FinOps. Ghosal’s account of the experiment is the source for its setup and results; the PlannerCritic repository provides project context and project-maintained field-test summaries.
The counts below describe blockers Ghosal reports in this sweep. They are not industry-wide rates, and the experiment has not been independently reproduced on the same corpus with the same blocker definitions.
Free tools Windows power users keep installed
One-click scans. No signup required.
The three recurring blocker families
| Blocker | Count in Ghosal’s sweep | What goes wrong | Example from the account |
|---|---|---|---|
| Unverified dependencies | 57 | A task assumes a condition that no earlier task establishes or checks. | Moving traffic to 100% without confirming stability at earlier traffic stages. |
| Unsafe sequencing | 46 | A task appears before a prerequisite that should come first. | Backfilling vectors before verifying index quality. |
| Weak rollback | 18 | The recovery step fails to address the state a high-impact change may have created. | Reverting dual-write mode without correcting possible inconsistencies. |
Ghosal says those three families account for 121 of the 132 concrete blockers he reports. The remaining blockers are outside those named categories; the article cautions that the three patterns may not cover domains the test did not include.
#1 Best Overall
1. Unverified dependencies
A plan can list all the right actions and still omit the evidence that makes the next action safe. If a cutover depends on a successful staged rollout, the plan needs an earlier task that verifies each required stage—not merely a task that performs the cutover.
2. Unsafe sequencing
Task order is a safety property, not a formatting detail. A later step may be reasonable only after its prerequisite has completed and its result has been checked. A plan that reverses that relationship can be unsafe even if every individual task sounds sensible.
Rank #2
3. Weak rollback
“Undo the change” is not always a complete recovery plan. A change may leave behind data inconsistencies or other state that survives a simple reversal. A rollback should address the failure state the change could produce, not just restore the previous mode or setting.
Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallOutdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchWhat Ghosal says the system measured
For the v0.2.1 sweep, Ghosal reports 170 goals at a total cost of $0.49. He also reports 13.86 seconds median latency for approved plans, 27.82 seconds for escalated plans, 2.58 mean blockers per goal, 58 escalation decisions per 100 goals, 1.4 mean LLM calls per goal, and a median of 1.0 revisions to resolution. These are measurements reported by the system’s author, not independently verified benchmarks.
In this experiment, Ghosal says using a larger model did not change the defect pattern. He also reports a critic trial in which label and evidence drift did not produce zero-blocker approvals on seeded defects. Those results describe these tests; they do not establish a general rule about model size or the reliability of LLM critics.
Why deterministic checks belong in the safety path
A model can help draft and critique a plan, but a hard invariant should not depend on a model noticing it every time. A deterministic gate can reject a plan when a required condition is absent from the plan’s verified history, or when the order violates a rule the system can encode. In Ghosal’s design, the critic can add findings but cannot suppress a blocker raised by a gate. As he puts it, “Deterministic gates catch what must be caught. The LLM critic is allowed to be unstable because it can only add findings, never suppress a gate blocker.”
The proposed precondition check asks whether every task’s preconditions are established by an earlier task. Ghosal estimates that this check would eliminate 64 of the 132 blockers, or 48%. He explicitly describes that figure as a projection, not a measured result after applying the fix. He also describes topological auto-repair for task ordering and oscillation detection for repeated plan cycles.
A practical way to apply the lesson
For teams building or evaluating automated change plans, the useful question is not simply whether a model can produce a coherent sequence. It is whether the system can verify the safety conditions that matter before action proceeds.
Best Value
- Make prerequisites explicit. For each task, state what must already be true and what evidence establishes it.
- Check dependencies and order in code. Reject plans whose required conditions are missing or whose task order violates an encoded dependency.
- Define recovery against possible resulting state. Specify how to detect and repair side effects, not only how to reverse the initiating change.
- Escalate unresolved plans. If a check cannot establish safety or a bounded repair loop cannot resolve the issue, require human review rather than treating uncertainty as approval.
- Keep approval separate from execution. The PlannerCritic repository says the software does not execute approved plans and does not guarantee their correctness. Approval by a planner is therefore not a substitute for operational controls.
What this experiment does—and does not—show
The result is a useful warning about a particular failure class: a plan can be convincing in language while missing a prerequisite, putting a step in the wrong order, or failing to restore a safe state. Ghosal’s sweep gives concrete examples and counts from one reported system and corpus. It does not establish how common these problems are across all AI planners, nor that the three named categories are exhaustive.
The engineering principle is narrower and more actionable: encode safety rules that can be stated precisely, make failures block or escalate, and use model judgment as an additional source of findings rather than the sole safety barrier.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.

