Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

If Claude Code reports a task as finished while your test suite is still failing, a Stop hook can make the test run part of the finish line. The approach described in a DEV Community post by elijahmanlockedin112, published September 26, 2026, is a short Python script that runs your verification command each time the agent tries to end a turn. If the command fails, the script returns the tail of the output to the agent with an instruction not to weaken or skip tests, so the agent keeps working instead of stopping. The design is the author’s. Treat its exact behavior as something to confirm against Anthropic’s current hooks reference for your Claude Code version before you rely on it.

How the gate works

The gate sits on the Stop event, which fires when Claude Code is about to end a turn. Each time that happens, the script follows the same sequence:

  1. It reads the hook’s JSON input from standard input.
  2. It finds the project directory from CLAUDE_PROJECT_DIR, falling back to the current directory.
  3. It reads a single verification command from .claude/verify.txt. If that file does not exist, the script exits successfully and the turn ends.
  4. It checks the consecutive_blocks counter. Once that counter reaches four, the script also exits successfully, so a stubborn failure cannot trap the session in an endless loop.
  5. It runs the command in the project directory with a 300-second subprocess timeout and captures the combined output.
  6. If the command exits with a nonzero status, the script prints the last 40 lines of output, adds a request to fix the failures without weakening or skipping tests, and exits with code 2. In the author’s design, code 2 is what blocks the stop and passes the message back to the agent.

The effect is that “done” becomes a claim the agent has to back up with a passing run, not just a sentence it writes after the last edit.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Setting it up

  1. From the project root, create the folders: mkdir -p .claude/hooks.
  2. Save the author’s script as .claude/hooks/verify-gate.py. Copy it from the original post, since the logic depends on the exact field names and exit behavior described there.
  3. Write your verification command into .claude/verify.txt as a single line. Run that command by hand first and confirm it passes on a healthy tree.
  4. Register the script under the Stop event in .claude/settings.json. The author’s example uses a hook timeout of 330 seconds, which is slightly longer than the script’s own 300-second subprocess limit so the script can finish cleanly before the hook is killed.

The registration follows the shape the author shows. Check the field names and the unit of timeout against the current hooks reference before you copy it:

{
  "hooks": {
    "Stop": [
      {
        "hooks": [
          {
            "type": "command",
            "command": "python3 .claude/hooks/verify-gate.py",
            "timeout": 330
          }
        ]
      }
    ]
  }
}

On Windows, the article says to use python in place of python3.

Choosing the verification command

The gate only checks what its command checks, so the command matters more than the script. The author’s guidance is to pick a noninteractive command that covers the feature you are building and finishes quickly enough to run after every turn. The author’s Node example chains two checks. The equivalents for other stacks are below; these are common tools and commands, not configurations the author tested.

Stack Example command for verify.txt What to confirm first
Node.js / TypeScript npm test && npx tsc --noEmit (the author’s example) Tests run without a watch mode or prompt
Python pytest -q pytest is installed in the project environment the hook uses
Go go test ./... Modules download without an interactive login step
Rust cargo test The build cache is warm enough that a run stays short

A command that exceeds the timeout, waits for input, or runs a full end-to-end suite on every turn will make the gate feel like a hang rather than a safeguard. Scope the command to the code you are changing where your test layout allows it.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Write the failing check first

The author recommends writing the test before the implementation and confirming that it fails while the feature is absent. This matters because a gate that passes on an empty project proves nothing. Run the verification command once before any code exists, see it fail for the expected reason, and only then let the agent build the feature. A test that was never red has not shown that it can catch anything.

Known limitations in the simple version

  • Runs with no changes. The author notes that the simple version can still run the command when no files changed, which adds time to turns that did no testable work.
  • Child processes after a timeout. When the command times out, child processes can be left running. Check for stray test runners after a timeout.
  • Missing counter field. The author says the consecutive_blocks field may not always arrive. If it is absent, the loop guard cannot count blocks, so a persistent failure may keep blocking.
  • Error policy. The author says that when an unexpected error occurs, the script should let the session continue rather than trap it. That favors availability over strictness, so an internal bug in the gate can silently remove the check.

The author describes an expanded version that addresses these cases. The original post is the place to find it, and this article does not cover that version’s code.

The trust boundary: who can change the check

The biggest weakness is not in the script’s logic. If the agent can edit the hook, the verification command, or the test assertions inside the working tree, it can make the gate pass by changing the check instead of the code. A commenter on the original post raised this point. Their suggested hardening stores the command and an expected-pass baseline at a path the agent’s permission tier can read but not write. That is a recommendation from a commenter, not a tested guarantee, and this article has not verified it.

Without that kind of separation, a practical safeguard is to review changes under .claude/ and to scan test files for newly added skips, deleted assertions, or loosened expectations before you accept a “done” result. A short git diff on those paths after each session shows whether the gate was edited along with the code.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Troubleshooting

  • The gate never runs. Confirm that .claude/settings.json is valid JSON, that the hook sits under Stop, and that the script path is correct relative to the project directory.
  • Python not found on Windows. Replace python3 with python in the registered command, as the author advises.
  • The turn stalls. The verification command is probably slow or interactive. Run it by hand, remove any watch flags, and shorten its scope.
  • The session ends with failing tests. The loop guard may have triggered after four blocks, or the counter field was missing. Read the last output the gate printed and fix the failure directly.
  • Test processes remain after a stop. A timed-out command can leave children running. Stop them manually before the next run.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What is and is not established

The author describes the behavior above and reports a failing test that the gate caught. That account is not independently reproduced here, and no repository or test transcript accompanies the post. The details that depend on Claude Code itself, including whether exit code 2 blocks the stop, whether stderr is returned to the agent, and which hook fields are always supplied, should be checked against Anthropic’s current hooks documentation for the version you run. Anthropic’s setup documentation covers installation and access, not these hook semantics.

The post also repeats a claim that teams see a two-to-three-times quality gain from verification-first workflows, attributed to a named engineer. No primary source for that figure is given, so it is not used as evidence here.

For a small project, the value of the gate comes from a fast, honest verification command that cannot be quietly edited by the agent. The script is a reasonable starting point. The integrity of the check is the part you have to design yourself.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.