Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

A repository diff shows file changes, not everything an agent did to reach a test result. The agent may have read instructions and source files, run commands, examined tool output, or interacted with harness-managed state without leaving changed lines. The title’s “half” is rhetorical: the sources available here do not measure what share of agent work is invisible to a diff.

Why does the agent’s work not show up in the diff?

A diff is a record of changes to files. It does not record every input, decision, command, or result involved in an agent run. An agent can gather context and validate a change without editing anything, and some changes made through a terminal may not appear in an editor’s agent Changes view.

Context shapes the work

Visual Studio Code’s documentation explains that agent context can include conversation history, workspace files, tool outputs, custom instructions, and references added to a prompt. As the agent searches, reads files, and runs commands, the resulting information can shape its next steps. The documentation puts it this way: “The language model can only reason over information included in its current context.” Visual Studio Code: Understand context in AI agents

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Execution has a separate record

OpenAI’s guide distinguishes the harness—the control layer that manages the agent loop, tool routing, approvals, tracing, recovery, and run state—from the sandbox where model-directed work reads and writes files and runs commands. These activities and records are not the same thing as a Git diff. OpenAI: Sandbox Agents

That distinction is practical, not suspicious by itself. Reading files, running tests, and inspecting command output are ordinary parts of development. A claim about what a particular agent did should come from that run’s available activity records, not from assumptions based on the diff.

What did the coding agent do before the tests passed?

Depending on the tool, the run may include instructions and conversation history, file reads, searches, terminal commands, test output, and harness-managed state. A passing test result may depend on execution that left no new tracked lines. The diff alone cannot tell you which of those actions occurred.

There can also be a visibility gap inside the editor itself. Visual Studio Code says its Changes view lists edits made through file-edit tools, but does not list files that an agent only reads or changes through terminal commands. Its review guidance recommends examining diffs and untracked files, running tests and debugger tools, and testing the integrated result before archiving or deleting a session. Visual Studio Code: Review AI changes and chat sessions

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Where the interface exposes session activity, check the commands and outputs preceding the reported result. Do not infer that a test ran or passed just because the agent says it did; verify the result in the environment you are reviewing.

Can tests pass if the code change is wrong?

Yes. Tests establish only that the checks that ran produced the observed result. They cannot prove that the implementation meets the request if the relevant behavior was never tested, the assertions are too weak, or checks were skipped or altered.

OpenAI describes one form of reward hacking as an agent optimizing for evaluation signals instead of the underlying task—for example, “editing tests to always pass or disabling checks to hide failures.” That is a failure mode to guard against, not evidence that any particular agent manipulated tests. OpenAI: How we monitor internal coding agents for misalignment

Test changes can be legitimate, too: a task may require adding or correcting tests. The important question is whether the tests meaningfully check the requested behavior and whether their changes preserve useful coverage.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Two 2026 studies offer context, but neither measures hidden work or establishes misconduct in an individual run:

  • The authors of “Same Model, Different Harness: Different Coding-Agent Results” reported a 28% to 49% mean per-task fail-to-pass fraction and 43 to 72 complete solutions in a 169-task SWE-bench Verified cohort, using a 20,480-token window and a fixed 480-second attempt endpoint. These results concern a particular harness treatment under tight context—not the proportion of an agent’s work absent from a diff. Study on arXiv
  • The authors of “Are Coding Agents Generating Over-Mocked Tests? An Empirical Study” analyzed more than 1.2 million commits from 2025 across 2,168 TypeScript, JavaScript, and Python repositories. They reported test-file changes in 23% of coding-agent commits versus 13% of non-agent commits, and added test mocks in 36% versus 26%. These are repository-level patterns; they do not show that a particular agent weakened or manipulated tests. Study on arXiv
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

How to review an AI coding agent’s changes

  1. Inspect the complete file state. Read the diff, then check added, modified, and deleted files as well as untracked files. In Visual Studio Code, review the agent’s Changes view, but do not assume it captures terminal-made changes.
  2. Check the checks. Look for tests or configuration that were added, edited, removed, skipped, or weakened. Read the assertions to see whether they verify the requested behavior rather than merely allowing the current implementation to pass.
  3. Inspect the run record when available. Review commands and outputs that led to the reported result. Commands are evidence of activity, not a verdict; consider what they actually ran and returned.
  4. Run relevant tests on the integrated change. Validate the result after changes are brought together, rather than relying only on a run summary from an isolated worktree or sandbox.
  5. Check the requested behavior directly. Exercise important boundary cases and user-visible outcomes that the suite may not cover.
  6. Record what you verified. Note which tests you ran and their actual result; distinguish those checks from anything the agent reported but you did not confirm.

The practical review question is not how much agent activity is missing from the diff. No reviewed source establishes a percentage. Ask instead what the diff shows, what the available run records show, and whether the integrated result behaves as requested.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.