Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

In the case Felixwang007 describes, yes. A code-review command received the word selftest as a positional argument, treated it as a scan path, found that no such path existed, and skipped it. The run reported a scope of zero files and zero lines, then printed a health score of 100/100, a “safe to merge” verdict, and exit code 0. The account comes from the author alone and has not been independently reproduced, but the failure pattern is worth understanding because it applies to any gate that treats “the process finished” as “the work was checked.”

What the author reports happened

The sequence, as the author lays it out in a write-up republished at World Programming on October 1, 2026, runs in five steps:

  1. The tool received the positional argument selftest.
  2. It interpreted that word as a path to scan.
  3. The path did not exist, so the scanner skipped it without raising an error.
  4. The run reported a scope of zero files and zero lines.
  5. The output showed 100/100 and “safe to merge,” and the process exited with code 0.

The author’s point is that the tool’s real self-test is invoked with --selftest. The malformed positional call therefore did not reach the self-test at all. It exposed a separate scan-mode behavior in which an unmatched input produced a clean-looking result instead of an error.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Two invocation conventions and one ambiguous word

A flag and a bare word look different to a human reader, but a parser that accepts both can map them onto very different code paths. The author’s audit of 34 packages found three conventions in use:

Self-test convention Packages in the author’s audit (of 34) Why it matters for a gate
--selftest flag 5 Explicit and unlikely to be mistaken for a path.
Positional selftest command 12 Indistinguishable from a path name when the scanner does not check whether the argument exists as a directory first.
No self-test 17 Nothing in the package can show that the tool detects a broken rule.

Counts are the author’s, from Felixwang007’s 2026 package audit. They describe that one audit, not agent tooling in general.

Why “nothing examined” can look like “nothing wrong”

An exit code of 0 tells a pipeline that the program did not crash. It does not tell the pipeline how much input was processed. A health score built from findings can also read as perfect when no findings are produced, whether because the rules ran cleanly or because no files reached them. The author’s central line puts the distinction plainly:

“Nothing wrong” and “nothing examined” are not the same result, and a tool that returns the same status for both cannot be part of a gate.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The author’s test for a bad report is equally direct: “If it says ‘safe,’ ‘passed,’ or ‘0 issues,’ you’ve found the bug.” That framing matters most in three places:

  • A CI step where a changed-file variable is empty, so the reviewer receives no files and the job still goes green.
  • An agent that passes the wrong parameter and then reports the change as reviewed.
  • A downstream step that trusts the exit code and never reads the scope line of the output.

The audit figures, as the author reports them

All numbers in this section are from Felixwang007’s 2026 audit and are not independently measured.

Flag-based self-tests

Five packages used a --selftest flag. The author reports their assertion counts as 30, 54, 16, 40, and 77.

The code-review tool’s self-test

The code-review tool’s --selftest is reported to have 54 assertions and 38 rules. In the author’s run, 23 of those rules fired on dirty samples, meaning deliberately flawed inputs designed to trigger them. The author does not state how many rules were checked against clean samples in the same run, so the 23 figure should not be read as a full pass or coverage rate.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Aggregate assertions

Across the 17 packages that had any self-test, the author gives a rounded total of about 700 assertions. That is the author’s own rounding, not a separately measured industry figure.

Paired positive and negative cases

The author argues that a checker needs both kinds of sample, because a rule that fires on everything and a rule that fires on nothing both pass a happy-path demo. The author’s examples come from a SQL inspector:

  • DROP TABLE should produce a finding; DROP TABLE IF EXISTS should not.
  • A phrase inside a string literal should not be mistaken for a missing WHERE clause.
  • SELECT * inside a comment should not be reported.
  • An environment variable reference should not be treated as a literal password.
  • A PL/pgSQL BEGIN ... END body should not be mistaken for an unclosed transaction.

These are the author’s examples. They are not independently tested behavior.

The author’s safeguards

The following are Felixwang007’s recommendations. The article does not cite a formal standard or independent validation for them.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  1. Report whether work was actually examined. Treat zero files, zero rules run, or zero tokens as a distinct non-success condition.
  2. Document one exact self-test invocation for each package, and have the harness read that contract rather than guess from source text.
  3. Demonstrate that a self-test can fail by intentionally breaking an assertion or rule.
  4. Include positive samples where a check must fire and negative samples where it must stay silent.
  5. Run the gate in the publish or deploy step and stop on failure, rather than trusting an earlier report.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

How much the evidence supports

This is one author’s package audit and one incident narrative. The exact behavior of the tool was not independently reproduced, and the audit figures are not representative of agent tooling generally. What the account does establish, in the author’s terms, is that an exit code and a score can both be clean while the scope of work is empty, and that a gate has to make that case visible. The tool’s own README-level caveat, quoted in the author’s write-up, makes the same boundary explicit:

static rules can only disprove, not prove — still verify permissions, concurrency and money precision by hand.

A passing static check is evidence that listed rules did not fire on the inspected input. It is not evidence that the change is correct, and it says nothing about files the tool never received.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.