What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
Sergey Petrukovich’s review of skillmem, a local memory tool for coding agents, found serious defects after the code had already been reviewed—and fixes sometimes introduced new ones. One correction matters most: although the original post suggested its two-clean-round stopping rule had been reached, Petrukovich later reported that the rule was not met during 26 additional rounds. His account is a project retrospective, not proof that any number of reviews or model pairing guarantees safer code.
Why did reviewed code need so many more rounds?
Petrukovich describes forty adversarial review rounds between skillmem releases 0.10 and 0.11.0. The project is local, SQLite-backed memory software for coding agents; its repository and README describe the project. Reviewers were expected to support findings with file and line details, severity, and reproducible command output. The point was not simply to ask for another general pass: the reviewers probed assumptions and attempted to reproduce concrete failures.
In the first 18 rounds, which Petrukovich characterizes as an audit, he reports six P1 findings. Examples included HTTP write ownership, permissions for public skills, shared body files, overly broad trust grants, and path traversal during export. The fixes themselves took three rounds of repair, an early indication that correcting a finding did not automatically make the change safe.
Rounds 19–23 focused on release rehearsal. Testing init against a copied configuration exposed duplicated hooks after a virtual environment was moved. Petrukovich also reports backup defects: backups could overwrite one another, were created with 0644 permissions near an OAuth-account file, or were not byte-exact. He counted eleven P2 findings in installer code that had appeared to work.
#1 Best Overall
In rounds 24–40, the review returned to core modules and found issues such as private record titles appearing in conflict messages and backlinks, visibility filtering happening after pagination, importing a symlink target outside a vault, an overly broad pack-removal command, a missing Windows scheduler environment, and secret redaction that was not idempotent. In that last case, repeated redaction could change content hashes and drop approval.
These are the author’s reported findings in one project; the retrospective does not independently reproduce each bug. The breadth is still instructive: review uncovered access-control, privacy, import, installer, and platform-specific defects—not only flaws in whichever module had received the most attention.
What changed when fixes caused regressions?
Petrukovich’s September 18 correction says that, after the original post, reviewers found two more ways approved rules could be neutralized, along with a Windows-specific issue. Across the next 26 rounds, the proposed rule of stopping after two consecutive rounds in which neither reviewer reproduced a P1 or P2 was never satisfied. About half of the later findings, he says, were regressions from earlier fixes.
Two recurring patterns explain why a narrow fix can leave a system exposed. A guard may be added to one caller rather than to the shared operation, leaving another route unprotected. Or a read followed by a write may leave a gap in which the state changes before the write takes effect. The response described in the retrospective was to put checks in the shared operation and make affected writes transactional.
Rank #3
The search fix that changed twice
Filtered search illustrates the trade-off between correctness and performance. One approach let hidden rows crowd out visible results; another made sorting slow. On skillmem’s 9,000-row database, Petrukovich reports that one implementation took 20 seconds per request. A later narrowed query took 52 ms unfiltered and 73–87 ms filtered. Those are measurements for his project and implementation, not general benchmarks.
Even after the performance work, a later iteration changed ranking in the master HTTP interface relative to the command-line interface. The example shows why a reviewer needs to check related entry points and expected behavior, not just whether the specific symptom that prompted a patch has disappeared.
What does the retrospective suggest about review practice?
Petrukovich recommends reviewers from different model lineages, but his account does not establish that a particular pairing is universally superior. The practical value of the workflow lies in making claims testable and fixes auditable:
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Fix the driver behind crashes, sound loss and screen glitches3Repair Windows errors before they cause bigger problems- Require each finding to identify the file and line, state severity, and include reproducible command output.
- Keep reviewers from changing code, so review findings remain distinct from implementation decisions.
- Check repository status after each round to catch unintended changes.
- Review every fix—including one-line changes—as a new change with its own failure modes. Petrukovich’s concise formulation is: “Every fix is a new round, one-liners especially.”
- Place security and correctness checks at shared mutation boundaries, and use transactions where a read-then-write gap matters.
These are practices grounded in one author’s experience, not a controlled comparison showing how much safer a codebase becomes. The reported round count, bug count, and measurements should be read as project-specific evidence.
Best Value
When should a team stop reviewing?
The original post proposed stopping after two consecutive rounds in which neither reviewer reproduced a P1 or P2. Its correction is the crucial qualification: when review continued across additional rounds and newly examined issues, the two-clean-round condition was never met. A quiet pair of rounds therefore did not demonstrate that all important defects had been found.
A stopping rule can still make a review more deliberate, but it should not be treated as a safety guarantee. Teams can record which code and call paths were reviewed, whether findings reproduce, whether fixes were re-reviewed, and whether a new module or platform issue has reopened the scope. If scope changes or a fix reveals a related path, the earlier clean rounds may no longer be meaningful for that expanded review.
The retrospective also records a feature-level decision: a proposed mem_archive feature accumulated 13 P1 findings over ten rounds, so Petrukovich removed it and made retirement an owner terminal command. He also describes an unattended review/fix/test loop that produced 415 tests rather than 346. Neither figure establishes a general return on investment; both describe choices and output in this project.
The original retrospective and its correction are available in Petrukovich’s DEV Community article. The project’s v0.11.1 release page and issue #5 provide additional project context; repository and release details can change over time.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

