Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

A better score in a “with skill” run does not prove the skill produced the improvement. In Driftproofhq’s September 20, 2026 evaluation, the largest apparent gain came in plugin runs whose traces reportedly showed that the skill was never invoked. Before attributing a result to a skill, check two things: Was the skill actually invoked? And did both arms sit at the ceiling?

The results are exploratory observations from a small number of cases, not general evidence about how effective agent skills are. The key lesson is methodological: an evaluation can test whether a model discovers a skill, whether it applies a skill it has already seen, or whether its output clears a grading threshold. Those are different questions.

What the evaluation compared

Driftproofhq’s September 20, 2026 article describes tests of three skills from the public addyosmani/agent-skills pack: code review and quality, git workflow and versioning, and documentation and ADRs. The author used one case for each skill and repeated each case with and without the skill. The report compares two evaluation methods, but explicitly treats them as measures of different things.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Plugin evaluation: can the model discover and invoke the skill?

In the built-in plugin evaluation, the skill is installed as a plugin. The model must discover and invoke it, then use it while working with tools in a workspace. Runs receive a pass or fail grade. This setup can measure activation as well as task performance.

Runner evaluation: can the model apply a skill it is given?

In the author’s runner, the skill text is placed directly into the model’s context, guaranteeing exposure. The runner then assigns a continuous score from zero to one over multiple draws. It assesses application given exposure; it does not measure whether the model would discover or invoke the skill on its own.

The author reports using claude-opus-5 as both target model and judge in both approaches. Because the methods differ in exposure, grading scale, and what they measure, their numerical results should not be compared as if they were interchangeable.

Was the skill actually invoked?

The most striking result concerned documentation and ADRs: Driftproofhq reports that all three plugin runs passed, compared with one of three runs without the plugin. But the author also reports that the Skill tool was not called in any of those three plugin runs, or in a supplementary run.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

That trace changes how the apparent gain should be read. A with-skill condition is not necessarily a skill-assisted condition: if the model never invokes the skill, the result cannot establish that the skill caused the difference. As Driftproofhq puts it, “If activation is not recorded anywhere, the difference is not attributable to the skill, whatever its size.” The reported difference may be worth investigating, but the traces do not support crediting it to skill activation.

For any evaluation that claims a skill helped, inspect the invocation record—not just the final score. Confirm that the skill was called in the relevant run and, where possible, that its content informed the work. If invocation is absent or unlogged, describe the comparison as a difference between conditions, not evidence of the skill’s effect.

Did both arms sit at the ceiling?

Pass/fail grading can conceal differences when both conditions already pass. In the code review and quality case, Driftproofhq reports three passes out of three both with and without the plugin: a zero pass-rate difference. Yet the continuous runner scored the case at 0.918 with the skill and 0.783 without it.

Those scores do not establish that the skill improved performance: they come from a different measurement method, and the report does not make a direct comparison between its tools. They do illustrate why a binary threshold can run out of room to show variation. As the author writes, “Both arms had cleared the pass threshold, so pass/fail had nothing left to report.” A zero difference in pass rate means the binary measure did not distinguish the conditions in that case; it does not, by itself, prove there was no effect.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What the reported numbers can—and cannot—show

All figures below are reported by Driftproofhq in 2026. They come from a small exploratory evaluation and are not independently verified benchmarks or broad estimates of skill effectiveness.

Case or check Reported result How to interpret it
Documentation and ADRs, plugin evaluation 3 of 3 runs passed with the plugin; 1 of 3 passed without it. The Skill tool was reportedly not called in the three plugin runs. The pass-rate difference is not attributable to skill activation based on the reported traces.
Code review and quality, plugin evaluation 3 of 3 runs passed in each condition; pass-rate difference: zero. Pass/fail did not distinguish the conditions in this case.
Code review and quality, continuous runner 0.918 with the skill and 0.783 without it. A separate, continuous scoring method; not directly comparable to plugin pass rates.
Tool-using sessions overall 18 of 18 sessions passed native grading and post-session verification: 9 with the plugin and 9 without. These checks did not ensure every claim was grounded in fixture evidence.
One ADR example It passed structural checks but asserted repository history that the fixture did not supply. Structural compliance and tool use do not guarantee factual grounding. This was one identified example, not a claim about all 18 sessions.

In the native evaluation, three runs per condition mean each individual run shifts that condition’s pass rate by 33 percentage points. The report also gives plus/minus values for the runner as sample standard deviations across draws; those describe observed spread and are not confidence intervals with a stated coverage probability.

Why the evidence remains exploratory

  • Only one case per skill: results describe those particular cases, not the skills across tasks or settings. The author says the findings are exploratory: “It makes them exploratory, which is what one case per skill can support.”
  • Few plugin runs: three runs per condition leave each pass/fail outcome influential, as the 33-percentage-point step illustrates.
  • Different measurement designs: one method includes discovery and invocation and grades pass/fail; the other guarantees exposure and produces continuous scores. Their numbers answer different questions.
  • Same model generated and judged: using claude-opus-5 for both roles creates a self-preference risk identified by the author.
  • Grounding needs its own check: a run can pass structural checks yet make claims that the supplied fixture does not substantiate, as in the single ADR example.

A practical checklist for reading skill evaluations

  1. Identify the question being tested. Is the evaluation testing discovery and invocation, application after guaranteed exposure, or a broader task outcome?
  2. Inspect invocation traces. Verify whether the skill was actually called in each run. A condition label alone does not establish activation.
  3. Check for a ceiling. If both conditions pass every run, binary grading cannot reveal quality differences above the threshold. Treat equal pass rates as an inconclusive measure of variation, not automatic proof of no effect.
  4. Read the scale and repeat count. Distinguish binary outcomes from continuous scores, and note how much a single run changes a rate. Treat descriptive spread as descriptive unless the analysis supports an inferential claim.
  5. Verify more than structure. Check whether claims are supported by the available evidence, not only whether the output has the right form or the workflow used tools.
  6. Limit the conclusion to the cases tested. Small samples and one case per skill do not support broad claims about effectiveness.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.