iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
The first run failed because the skill gave a verdict when it had no evidence to support one. The fix was not a longer prompt. It was a rule that treats a missing sample as a reason to defer the decision and names the sample that would settle it. In a September 23, 2026 Dev.to article, Vishal Habib reports that the revised skill then passed its gates. These are the author’s own results from his setup. They have not been independently reproduced, so read them as one developer’s documented evaluation, not as a general measure of Claude Code skills.
What the author tested
Habib says he built three Claude Code skills for AI product managers and published the evaluation suite on GitHub, including the runs that failed. The sequence he describes matters more than any single number. He wrote down the pass criteria before running anything, so the criteria could not be adjusted to fit the outcome.
One of the three skills, /build-or-not, is meant to assess a feature idea against real examples. Its job is to help a product manager decide whether to build something, based on evidence rather than instinct.
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Fix the driver behind crashes, sound loss and screen glitches3Repair Windows errors before they cause bigger problemsWhy the first run scored 0.00
In the first test, the skill had no evidence to work with and no research tools were available. It still returned “don’t build.” According to the article, the verdict came from the model’s recalled market knowledge. The author reports a first-run score of 0.00 for this test.
#1 Best Overall
The cause, as he describes it, was a gap in the skill’s instructions: nothing told it what to do when no sample existed. A skill that is asked for a decision will produce one unless it is told that declining is acceptable. That is the failure the eval was designed to expose.
The fix: no sample, no decision
The author added a rule he calls “no sample, no decision.” Its effect is that the skill’s accepted outcomes now include “can’t decide yet,” and when it chooses that outcome it must name the sample that would resolve the question. The steps the skill now follows, as the article describes them, are:
Rank #2
- Check whether real examples for the feature idea are present in the input.
- If none are present, do not issue a build or no-build verdict.
- Return “can’t decide yet” and state the specific sample needed to decide.
- If a sample is present, assess the idea against it and give a verdict with its reasoning.
After this change, the author reports that the next run passed the gates. He puts the underlying principle in one line: “A bar set after the numbers can’t fail.”
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
What the eval measured
The reported evaluation covered eight cases, with three runs per case, on a single model. The article’s comparison table shows the skills performing better than plain Claude on several behaviors:
Rank #3
- stating the decision bar before deciding
- refusing to decide when there is no evidence
- planning a rollback trigger
- distinguishing a reasoned decline from an evidence gap
- reporting two separate coverage numbers
For four other cases, the article reports that plain Claude performed just as well as the skill-enabled version. Habib calls the exercise a check of key behaviors rather than a benchmark, and that framing is the right one for a sample this size.
What it cost
The article reports approximately $2 per full run. That figure belongs to the author’s setup, including his prompts, cases, and model. It is not a published Claude Code price, and it should not be used to estimate what another team would spend on a different suite.
Setting up your own pass bar
- Write the criteria first. Record what a passing run must show before you run anything, and keep that record unchanged afterward.
- Include a no-evidence case. Test whether the skill still produces a confident answer when the input is missing the material it needs.
- Allow a deferral outcome. If “can’t decide yet” is not a valid output, the skill will be graded on guesses.
- Keep failed runs. The first failure in this account is the most informative result in it.
- Record where plain Claude matches the skill. A skill that adds nothing on some cases is useful information too.
Checking activation and output quality separately
A GitHub-hosted copy of Claude Code skills documentation recommends evaluating two things separately: whether the skill activates when it should, and whether its output meets expectations. It suggests realistic prompts run in fresh sessions, with the skill enabled and disabled, so each case is compared against a clean baseline.
The same documentation copy describes a claude plugin eval command that runs plugin-on and plugin-off cases in isolated sessions and applies graders. Command names, flags, and installation steps change between versions. Confirm them against Anthropic’s current official Claude Code documentation before relying on them, because the copy’s currency against the live docs was not established here.
Best Value
Limits of the evidence
- The results come from one developer, one model, eight cases, and three runs per case.
- The author’s repository and the reported outcomes have not been independently rerun.
- The $2 cost applies only to the author’s setup.
- The comparison covers the behaviors the author chose to check, so it does not show overall skill quality.
What the account does establish is a method: fix the pass bar before testing, keep the failure, and let a skill say “can’t decide yet” when the evidence is missing.
Quick Recap
“
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

