Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

An open-source Claude Code project called pr-proof reports that its pull-request comment validator removed 34% of noise issues from CodeRabbit’s reviews in a benchmark—while retaining 72 of 77 labeled real bugs. Those are the project’s results on 50 benchmark pull requests, not a guarantee that it will improve every team’s reviews.

What pr-proof does

pr-proof is a public Apache-2.0 repository by TanayK07 containing three Claude Code skills for checking pull-request review comments against the code. The central idea is to treat each review comment as a claim: trace the relevant execution path, inspect callers, and check library behavior before deciding whether the finding holds up.

The project description puts it this way: “Every review comment has to prove itself before you see it.” In practice, the three skills have distinct jobs:

  • pr-comment-validation: Examines comments and labels each one valid, partly valid, wrong, or style, with code evidence. It does not change files.
  • pr-validation: Checks out a pull request in a worktree, validates its comments, shows the verdicts, can apply fixes you approve, and replies on review threads.
  • pr-review: Generates a new review, then has independent subagents try to disprove its findings before posting. It can also draft the review to a file.

These are not three names for the same workflow. The first validates comments without modifying anything; the second can carry approved fixes through a PR workflow; the third generates a review rather than filtering an existing one.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What the CodeRabbit benchmark found

The repository reports a run on 50 real pull requests from Code Review Bench. The benchmark includes PRs from Sentry, Grafana, Keycloak, Discourse, and Cal.com, with human-written “golden comments” used as labels. The repository does not state the benchmark year.

Measure CodeRabbit comments before filtering After pr-proof filtering
Issues 300 219
Precision 25.7% 32.9%
Recall 56.2% 52.6%
F1 35.2% 40.4%

In the same reported run, the validator retained 72 of 77 labeled real bugs (93.5%) and removed 76 of 223 issues classified as noise (34%). The F1 score increased by 5.2 percentage points, with a reported 95% confidence interval of +1.9 to +8.3 points. The project says the filter received each benchmark comment’s extracted text, file, and line along with checked-out code; it did not see the labels, and it was scored against the benchmark’s published labels. The README says both the published benchmark results and pr-proof run used Claude Opus 4.5 as judge.

These figures describe comment filtering, not a finding that CodeRabbit is generally unreliable. They also do not establish what every team will see on current private repositories, different languages, or a different Claude Code setup.

Why the benchmark is not a production guarantee

The repository lists several limits that matter when interpreting the result:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • The labels may be incomplete. The “golden” issue lists may omit real problems, so something scored as noise could be a valid but unlisted issue. That can make measured precision understate quality.
  • Public-code exposure is possible. The evaluated PRs are public and older than the models, so training-data leakage cannot be ruled out.
  • The evaluation environment was isolated. Runs were headless Claude Code sessions without user settings, hooks, MCP servers, plugins, web access, gh, or curl; they also could not read the original PR discussions.
  • Results varied across runs. The README reports two identical drafting runs with F1 scores of 33.5% and 28.2%. Its bootstrap confidence intervals are calculated over 50 PRs, so they do not remove uncertainty about behavior in a different set of repositories.

Taken together, the results support testing the validator as a filter on existing review comments. They do not establish a universal reduction in review noise or predict the outcome on a particular codebase.

The separate PR-review skill has a different result

Do not transfer the CodeRabbit filtering result to pr-review. For that separate skill, the README reports F1 of 29.8% (95% confidence interval 24.5–35.3%), compared with 29.1% (25.5–33.1%) for plain Claude Code Opus 5.5. The difference was +0.7 points, with a confidence interval of −3.4 to +4.8 points. The project describes the result as statistically level with plain Claude Code: pr-review wrote fewer comments and was more precise, but found fewer bugs.

The README’s own interpretation is that “The standalone validator is where pr-proof clearly earns its place, so that’s the headline above.” That is the project author’s reading of these benchmark results, not an independent assessment.

How to install pr-proof

The README lists Claude Code and an authenticated gh CLI as prerequisites. It offers a plugin installation route or manual copying of the skill folders.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Install as a Claude Code plugin

  1. In Claude Code, run /plugin marketplace add TanayK07/pr-proof.
  2. Then run /plugin install pr-proof@pr-proof.

Copy the skills manually

Alternatively, copy the folders under the repository’s skills/ directory into ~/.claude/skills/, following the project’s README.

Anthropic describes skills as instructions stored in a SKILL.md file: “Skills extend what Claude can do. Create a SKILL.md file with instructions, and Claude adds it to its toolkit.” Its Claude Code skills documentation explains that skills can load when relevant or be invoked with /skill-name, and that teams can share them in a project or distribute them through a plugin.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

How to invoke the three workflows

The repository’s example prompts show the intended distinction. Use a question about validity when you only want to assess comments, a request to handle a PR’s comments when you want the broader validation-and-fix workflow, and a review request when you want Claude Code to generate a review:

  • are these PR comments valid?
  • handle the review comments on PR #123
  • review PR #123

Choose based on whether you are validating an existing review or asking for a new one; the benchmark evidence for noise reduction applies to the former.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.