A scan of 45 human review comments across 15 merged Ruff pull requests produced two possible “unwritten review rules.” Neither was ready to become policy: the stronger-looking pattern appeared twice in the same pull request, while the other was too vague to enforce. The useful result was not a proven map of Ruff’s conventions, but a practical warning about how easily a small sample can mistake one review conversation for a recurring rule.
What the Ruff scan found
In a post published September 20, 2026, Ofer’s Instinct Bot describes running PR Rulebook over 15 merged pull requests in astral-sh/ruff. After excluding bot comments, the run contained 45 inline human review comments. The tool grouped some comments into two candidate rules. These counts and scores are the author’s account of that run, not independently verified measurements.
The plausible-looking pattern: mention async when it explains a diagnostic
One cluster suggested including the async keyword in a diagnostic annotation when it explains why the diagnostic fires. PR Rulebook assigned the cluster an 82% confidence score and found two accepted-change signals. But both comments came from the same pull request. They belonged to one review conversation, not evidence that the guidance recurred across separate PRs.
The weaker pattern: improve an error message
A second cluster combined two comments about quoting or improving an error message. Only one comment had an accepted-change signal, and the author judged the proposed rule too vague to enforce. Its reported score was 68%. Neither score should be read as a statistical probability that a rule is true: the author describes confidence scores as a way to rank candidates for inspection.
#1 Best Overall
Why repeated comments can mislead
Two comments can look like repetition without showing a team-wide convention. When both occur in one PR, they may reflect a single discussion, a reviewer’s follow-up, or the context of that change. Counting comments alone obscures that dependency. For a claim about recurring practice, the more useful unit is the distinct pull request: does the pattern appear in separate review conversations?
The author also identifies a text-processing trap. GitHub’s fenced suggestion blocks can contain shared placeholder text; lexical overlap in those blocks may cause unrelated comments to be grouped together. Similarity between words is not enough to establish that comments express the same actionable rule.
Rank #2
What the author says changed in PR Rulebook
According to the post, the revised process requires evidence from at least two distinct PRs before treating a pattern as a candidate rule. It also removes fenced suggestions before clustering, preserves identifiers inside inline code, canonicalizes a small set of review concepts, and compares normalized terms using cosine similarity. The author says a regression test rejects repeated comments confined to one PR. These are implementation claims from the tool’s author; the code and test were not independently inspected for this account.
| Check | Why it matters |
|---|---|
| Evidence spans distinct PRs | Helps distinguish recurring feedback from repetition within one conversation. |
| Accepted-change signals are present | Shows whether the author’s described workflow connected a comment to a code change, though it does not by itself prove a general convention. |
| Rule wording is specific | A candidate such as “mention async when it explains the diagnostic” is more actionable than “improve the error message.” |
| Suggestion blocks and inline identifiers are handled deliberately | Reduces misleading text overlap while retaining meaningful code terms, as described by the author. |
| A person approves the candidate | Keeps a similarity score from becoming an automatic policy decision. |
Candidate rules are prompts for review, not policy
PR Rulebook is described as a local TypeScript CLI that scans merged pull requests for recurring human review comments followed by code changes. It emits ranked candidate rules with evidence links and confidence scores, with output formats named for Cursor .mdc, Claude Code markdown, CodeRabbit YAML, and JSON. The author says a human approves each rule before it reaches coding agents.
Recommended Free Tools
That boundary matters. A candidate generator can help surface patterns worth checking, but this example does not establish the tool’s accuracy, a general threshold that works across repositories, or the prevalence of any particular review practice. The post reports no controlled benchmark, baseline, or independent precision-and-recall results. As the author puts it: “Confidence scores rank what a human should inspect. They do not make weak evidence true.” — Ofer’s Instinct Bot, original post.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Availability and limits of the account
The post describes PR Rulebook as unpublished on npm at the time of publication and invites five public repositories with active human PR review to join a pilot. Its from-source instructions specify Node 20 or later, installation and build commands, a read-only GitHub token, and a Markdown output file. The author says repository contents and comments travel from GitHub to the user’s machine and that PR Rulebook has no server. Availability and architecture here are the author’s description, not independently verified facts.
This is a builder’s self-reported case study. The underlying 15 pull requests, 45 comments, implementation, and regression test were not independently inspected. The scan is therefore useful as an example of a failure mode and a proposed safeguard—not proof of Ruff’s team conventions or of how well review-comment mining works generally.
Quick Recap
Best Value
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.
Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Fix the driver behind crashes, sound loss and screen glitches3Clear out junk files and repair common Windows errors

