Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

In a reported test of 170 change-planning goals, Debashish Ghosal says his AI planning system repeatedly ran into three structural problems: unverified prerequisites, unsafe task order, and weak rollback plans. His practical takeaway is to put hard safety rules in deterministic checks and send unresolved plans to a human—not to assume that a plausible plan is a safe one.

What the 170-goal test involved

Ghosal describes PlannerCritic as a workflow in which one large language model drafts a structured plan, deterministic gates check it against hard rules, and a second model critiques plans that pass those gates. A bounded revision loop can make further changes or escalate the plan for human review. The reported sweep covered 40 domains, including identity management, multi-agent operations, site reliability engineering, supply-chain policy, and FinOps. Ghosal’s account of the experiment is the source for its setup and results; the PlannerCritic repository provides project context and project-maintained field-test summaries.

The counts below describe blockers Ghosal reports in this sweep. They are not industry-wide rates, and the experiment has not been independently reproduced on the same corpus with the same blocker definitions.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The three recurring blocker families

Blocker Count in Ghosal’s sweep What goes wrong Example from the account
Unverified dependencies 57 A task assumes a condition that no earlier task establishes or checks. Moving traffic to 100% without confirming stability at earlier traffic stages.
Unsafe sequencing 46 A task appears before a prerequisite that should come first. Backfilling vectors before verifying index quality.
Weak rollback 18 The recovery step fails to address the state a high-impact change may have created. Reverting dual-write mode without correcting possible inconsistencies.

Ghosal says those three families account for 121 of the 132 concrete blockers he reports. The remaining blockers are outside those named categories; the article cautions that the three patterns may not cover domains the test did not include.

1. Unverified dependencies

A plan can list all the right actions and still omit the evidence that makes the next action safe. If a cutover depends on a successful staged rollout, the plan needs an earlier task that verifies each required stage—not merely a task that performs the cutover.

2. Unsafe sequencing

Task order is a safety property, not a formatting detail. A later step may be reasonable only after its prerequisite has completed and its result has been checked. A plan that reverses that relationship can be unsafe even if every individual task sounds sensible.

3. Weak rollback

“Undo the change” is not always a complete recovery plan. A change may leave behind data inconsistencies or other state that survives a simple reversal. A rollback should address the failure state the change could produce, not just restore the previous mode or setting.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What Ghosal says the system measured

For the v0.2.1 sweep, Ghosal reports 170 goals at a total cost of $0.49. He also reports 13.86 seconds median latency for approved plans, 27.82 seconds for escalated plans, 2.58 mean blockers per goal, 58 escalation decisions per 100 goals, 1.4 mean LLM calls per goal, and a median of 1.0 revisions to resolution. These are measurements reported by the system’s author, not independently verified benchmarks.

In this experiment, Ghosal says using a larger model did not change the defect pattern. He also reports a critic trial in which label and evidence drift did not produce zero-blocker approvals on seeded defects. Those results describe these tests; they do not establish a general rule about model size or the reliability of LLM critics.

Why deterministic checks belong in the safety path

A model can help draft and critique a plan, but a hard invariant should not depend on a model noticing it every time. A deterministic gate can reject a plan when a required condition is absent from the plan’s verified history, or when the order violates a rule the system can encode. In Ghosal’s design, the critic can add findings but cannot suppress a blocker raised by a gate. As he puts it, “Deterministic gates catch what must be caught. The LLM critic is allowed to be unstable because it can only add findings, never suppress a gate blocker.”

The proposed precondition check asks whether every task’s preconditions are established by an earlier task. Ghosal estimates that this check would eliminate 64 of the 132 blockers, or 48%. He explicitly describes that figure as a projection, not a measured result after applying the fix. He also describes topological auto-repair for task ordering and oscillation detection for repeated plan cycles.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

A practical way to apply the lesson

For teams building or evaluating automated change plans, the useful question is not simply whether a model can produce a coherent sequence. It is whether the system can verify the safety conditions that matter before action proceeds.

  1. Make prerequisites explicit. For each task, state what must already be true and what evidence establishes it.
  2. Check dependencies and order in code. Reject plans whose required conditions are missing or whose task order violates an encoded dependency.
  3. Define recovery against possible resulting state. Specify how to detect and repair side effects, not only how to reverse the initiating change.
  4. Escalate unresolved plans. If a check cannot establish safety or a bounded repair loop cannot resolve the issue, require human review rather than treating uncertainty as approval.
  5. Keep approval separate from execution. The PlannerCritic repository says the software does not execute approved plans and does not guarantee their correctness. Approval by a planner is therefore not a substitute for operational controls.

What this experiment does—and does not—show

The result is a useful warning about a particular failure class: a plan can be convincing in language while missing a prerequisite, putting a step in the wrong order, or failing to restore a safe state. Ghosal’s sweep gives concrete examples and counts from one reported system and corpus. It does not establish how common these problems are across all AI planners, nor that the three named categories are exhaustive.

The engineering principle is narrower and more actionable: encode safety rules that can be stated precisely, make failures block or escalate, and use model judgment as an additional source of findings rather than the sole safety barrier.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.