An open-source Claude Code project called pr-proof reports that its pull-request comment validator removed 34% of noise issues from CodeRabbit’s reviews in a benchmark—while retaining 72 of 77 labeled real bugs. Those are the project’s results on 50 benchmark pull requests, not a guarantee that it will improve every team’s reviews.
What pr-proof does
pr-proof is a public Apache-2.0 repository by TanayK07 containing three Claude Code skills for checking pull-request review comments against the code. The central idea is to treat each review comment as a claim: trace the relevant execution path, inspect callers, and check library behavior before deciding whether the finding holds up.
The project description puts it this way: “Every review comment has to prove itself before you see it.” In practice, the three skills have distinct jobs:
pr-comment-validation: Examines comments and labels each one valid, partly valid, wrong, or style, with code evidence. It does not change files.pr-validation: Checks out a pull request in a worktree, validates its comments, shows the verdicts, can apply fixes you approve, and replies on review threads.pr-review: Generates a new review, then has independent subagents try to disprove its findings before posting. It can also draft the review to a file.
These are not three names for the same workflow. The first validates comments without modifying anything; the second can carry approved fixes through a PR workflow; the third generates a review rather than filtering an existing one.
Recommended Free Tools
#1 Best Overall
What the CodeRabbit benchmark found
The repository reports a run on 50 real pull requests from Code Review Bench. The benchmark includes PRs from Sentry, Grafana, Keycloak, Discourse, and Cal.com, with human-written “golden comments” used as labels. The repository does not state the benchmark year.
| Measure | CodeRabbit comments before filtering | After pr-proof filtering |
|---|---|---|
| Issues | 300 | 219 |
| Precision | 25.7% | 32.9% |
| Recall | 56.2% | 52.6% |
| F1 | 35.2% | 40.4% |
In the same reported run, the validator retained 72 of 77 labeled real bugs (93.5%) and removed 76 of 223 issues classified as noise (34%). The F1 score increased by 5.2 percentage points, with a reported 95% confidence interval of +1.9 to +8.3 points. The project says the filter received each benchmark comment’s extracted text, file, and line along with checked-out code; it did not see the labels, and it was scored against the benchmark’s published labels. The README says both the published benchmark results and pr-proof run used Claude Opus 4.5 as judge.
Rank #2
These figures describe comment filtering, not a finding that CodeRabbit is generally unreliable. They also do not establish what every team will see on current private repositories, different languages, or a different Claude Code setup.
Why the benchmark is not a production guarantee
The repository lists several limits that matter when interpreting the result:
Rank #3
- The labels may be incomplete. The “golden” issue lists may omit real problems, so something scored as noise could be a valid but unlisted issue. That can make measured precision understate quality.
- Public-code exposure is possible. The evaluated PRs are public and older than the models, so training-data leakage cannot be ruled out.
- The evaluation environment was isolated. Runs were headless Claude Code sessions without user settings, hooks, MCP servers, plugins, web access,
gh, orcurl; they also could not read the original PR discussions. - Results varied across runs. The README reports two identical drafting runs with F1 scores of 33.5% and 28.2%. Its bootstrap confidence intervals are calculated over 50 PRs, so they do not remove uncertainty about behavior in a different set of repositories.
Taken together, the results support testing the validator as a filter on existing review comments. They do not establish a universal reduction in review noise or predict the outcome on a particular codebase.
The separate PR-review skill has a different result
Do not transfer the CodeRabbit filtering result to pr-review. For that separate skill, the README reports F1 of 29.8% (95% confidence interval 24.5–35.3%), compared with 29.1% (25.5–33.1%) for plain Claude Code Opus 5.5. The difference was +0.7 points, with a confidence interval of −3.4 to +4.8 points. The project describes the result as statistically level with plain Claude Code: pr-review wrote fewer comments and was more precise, but found fewer bugs.
Rank #4
The README’s own interpretation is that “The standalone validator is where pr-proof clearly earns its place, so that’s the headline above.” That is the project author’s reading of these benchmark results, not an independent assessment.
How to install pr-proof
The README lists Claude Code and an authenticated gh CLI as prerequisites. It offers a plugin installation route or manual copying of the skill folders.
Best Value
Install as a Claude Code plugin
- In Claude Code, run
/plugin marketplace add TanayK07/pr-proof. - Then run
/plugin install pr-proof@pr-proof.
Copy the skills manually
Alternatively, copy the folders under the repository’s skills/ directory into ~/.claude/skills/, following the project’s README.
Anthropic describes skills as instructions stored in a SKILL.md file: “Skills extend what Claude can do. Create a SKILL.md file with instructions, and Claude adds it to its toolkit.” Its Claude Code skills documentation explains that skills can load when relevant or be invoked with /skill-name, and that teams can share them in a project or distribute them through a plugin.
How to invoke the three workflows
The repository’s example prompts show the intended distinction. Use a question about validity when you only want to assess comments, a request to handle a PR’s comments when you want the broader validation-and-fix workflow, and a review request when you want Claude Code to generate a review:
are these PR comments valid?handle the review comments on PR #123review PR #123
Choose based on whether you are validating an existing review or asking for a new one; the benchmark evidence for noise reduction applies to the former.
The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

