Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Repair Windows errors before they cause bigger problemsFix Now →iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
A ticket marked done proves that a workflow recorded completion; it does not prove that the promised result exists. In one two-week review reported by Agent Diary, 31 tickets were closed as done, but only 9 had a live post the author could open with a URL. That is one author’s account—not an industry success rate—but it exposes a useful reliability rule: close work against verifiable evidence of the artifact, not a green checkmark alone.
What the 31-to-9 gap does—and does not—show
Agent Diary’s account describes a review of its own closed tickets: 31 were marked done, while 9 had a live post accessible by URL. The figures are anecdotal and should not be read as a prediction of how often AI agents publish successfully elsewhere. The important point is the mismatch between workflow state and observable outcome.
The author identified three ways a green status overstated success: a submission had been saved as a draft; signup appeared to work, but the publishing attempt redirected to a login wall; and a stale cache row led a report to count a post that was not actually present. Each failure can leave a completion record looking convincing while the user-facing result is missing or inaccessible.
The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Why a status is not proof of publication
A status field answers a workflow question: where does the system think this ticket is? It does not necessarily answer whether the external action succeeded. Jira Work Management Cloud documentation distinguishes a status, which represents the current state, from a resolution, which represents the final state of the work. Atlassian’s documentation says, “The resolution closes the issue and represents the final state of the piece of work.” Its examples include resolutions such as done, published, and rejected, and its editorial workflow illustrates draft, review, and published states. See Atlassian’s workflow documentation.
#1 Best Overall
These distinctions are configurable workflow concepts, not a guarantee that a page is live. A ticket can transition to a terminal status even if the publication tool created only a draft, encountered an authentication wall, or relied on outdated reporting data. A reliable workflow therefore needs an acceptance check tied to the claim the ticket makes.
Make the completion check match the promised result
Agent Diary proposes making the terminal transition depend on evidence that the artifact exists. Its examples include a URL returning HTTP 200, a file existing on disk, or an account successfully logging in. These are useful starting points, but the check must be strong enough for the ticket’s actual promise: a successful HTTP response alone does not establish that the intended content is visible to the intended audience, complete, or persistent.
For a published post
- Open or fetch the expected URL, not merely a generic success page.
- Check that the response contains the expected post or another reliable identifier for it.
- Verify the visibility promised by the ticket. If it promises a public post, check without relying on the publishing session’s authenticated view.
- Record the URL and the result of the check so an operator can inspect the evidence later.
For a file or account task
- For a file, check that the expected path exists and, where integrity matters, that its contents match the expected output.
- For an account task, verify the authenticated action the ticket promised—not just that signup returned a success message.
The extra checks are practical applications of the artifact-gate idea: test the outcome the user asked for, rather than treating a tool’s success report as the outcome itself.
Do these 3 things before closing this tab:
1Scan for outdated or missing drivers - takes under a minute2Repair Windows errors before they cause bigger problems3Fix the driver behind crashes, sound loss and screen glitchesSeparate in-progress work from verified completion
Use workflow states to represent meaningful stages such as draft, review, and published, then make the final transition conditional on the publication check. This gives a reader or operator a more useful signal than a single broad “done” state. If the workflow cannot enforce the check, keep the state provisional until a person or separate monitor verifies the artifact.
Rank #3
For multi-agent systems, also be explicit about where verification evidence is stored and who can see it. Cloudflare’s Agents documentation describes one product-specific model: workflows are tracked in the originating agent’s database, while workflows started by a child agent are tracked in that sub-agent’s storage. A parent that needs a combined view must aggregate child information explicitly. This is an implementation example, not a universal rule for agent frameworks. See Cloudflare’s workflow documentation.
Keep evidence with the ticket
A practical audit trail can include the expected artifact, the check performed, its result, and a timestamp. For publication, that might be the destination URL plus confirmation that the expected content was visible. Retaining this evidence makes it easier to distinguish a genuine completion from a stale report, a draft, or a session problem when someone later asks why a ticket was closed.
Agent Diary says that after its proposed change, the number it trusted was zero tickets closed without an artifact. That is the author’s described metric after changing the process, not an independently audited result. The defensible takeaway is narrower: a workflow should not mark a promised outcome complete until it has evidence appropriate to that outcome.
Where broader agent reliability research fits
A September 29, 2026 preprint by Rohith Reddy Bellibatlu, Zichong Wang, and Wenbin Zhang audited 34 mutating tools across four benchmarks and reported seven tool defects and one evaluator property at pinned commits. The work concerns benchmark tool behavior, not production publishing or Agent Diary’s ticket tally; it does not establish a publishing success rate. It does, however, examine the general reliability problem of a tool reporting success while the environment state differs from the change it was meant to make. Read the preprint.
Quick Recap
Best Value
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

