iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
Rcids built jev-tools, a Claude Code plugin that uses a hosted model for small, structured decisions rather than asking it to write prose. The experiment suggests that explicit probabilities and code-controlled thresholds can make agent decisions easier to route—but its live tests were small, sometimes inconsistent, and written by the same person who built the checks. It is an early engineering report, not an independent benchmark.
What the plugin does—and what “yes/no” means here
Rcids describes OpenJev, served by Codiv, as a model that does not generate prose for this plugin. A request contains a state—either a string or JSON—plus typed questions. Depending on the question, the response can include a yes probability, a choice with probabilities, or an expected score with probabilities for each level. The plugin then applies ordinary code and thresholds to decide what to do.
That distinction matters: a typed response is easier for software to consume than free-form text, but it does not establish that the probabilities are calibrated or that a decision is correct. Codiv’s characterization of calibration was not independently verified in Rcids’s account.
Rcids wrote the plugin for Python 3.10 or later using the standard library. Its components cover several points in a coding workflow:
#1 Best Overall
- PreToolUse rule hook: On Edit and Write operations, it reads rules from
CLAUDE.mdorAGENTS.md, assesses whether a proposed edit breaks one, asks a stricter second question when a violation is suspected, and can block the edit in active mode. - Opt-in skill picker: It chooses an installed skill for a prompt, then checks whether the choice is appropriate.
- Review precheck: It asks seven yes/no questions about a git diff, including whether it contains secrets, dependency changes, authentication or schema changes, weakened tests, swallowed errors, or risky logic. The result routes the change to a fast or full review.
- Rule calibration: It replays recent commits against rules and labels each rule decisive, noisy, weak, or quiet before enforcement.
- Other tools:
find-filessearches for files in two stages;browser-navchooses a next click from interactive page elements; andstatuschecks whether installation wiring is alive.
The design choices are pragmatic: keep thresholds in code, decide failure behavior separately for each feature, start in shadow mode, and share a standard-library API client. Rule and skill hooks fail open during an outage; review precheck fails safe by routing to a full review. In shadow mode, hooks record what they would do without blocking or injecting a skill.
What broke in live use
Offline tests with mocks did not expose every issue. Rcids reports that live API calls and Claude Code sessions uncovered several, with fixes that illustrate where this design can be brittle.
API requests were rejected
Calls initially returned 403 responses because Codiv’s edge rejected Python’s default User-Agent. Rcids changed the client to send its own. A passing mock test had not caught the behavior of the live service boundary.
Recommended Free Tools
A clean edit looked like a rule violation
An edit that read its host from configuration received a 0.91 score against a rule against hardcoding API hosts. Rcids added a stricter second question; in this case, the second check vetoed the same edit. This is one example of a second look reducing a false positive, not evidence that the approach will resolve false positives generally.
Rank #2
Browser navigation misread a completed goal
The navigator classified a page where the goal was already met as blocked. Across identical calls, the probability that the goal was met varied around the decision bar: one call returned 0.93, while another fell below 0.8. Rcids added a rule that combines “no remaining clicks” with a likely-met goal. The episode shows why a single score near a threshold can produce unstable routing.
Skill selection needed confirmation
The picker could inject a skill based on a shallow match. Rcids added a second-stage confirmation that uses the full skill description. The additional check makes the decision less dependent on a brief match, but it also adds another model call.
File discovery was inconsistent
In one evaluation, file discovery scored 3 of 4 and then 0 of 4 after a rewrite. Rcids responded by evaluating six variants rather than trusting one run. The later comparison with keyword counting, described below, still found no demonstrated advantage.
Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Scan for outdated or missing drivers - takes under a minute3Clear out junk files and repair common Windows errorsWhat the September 2026 measurements show
Rcids says these live tests ran against hosted OpenJev in September 2026 and calls them smoke tests, not benchmarks. The samples are small, noisy, and in part authored by the same person who built the system.
Rank #3
| Feature or measure | Rcids’s reported result | What the result does—and does not—show |
|---|---|---|
| Rule enforcer | 4 of 4 planted violations blocked and 4 of 4 clean edits allowed after the second-look check. Rcids also reports correct blocking and allowing in real headless Claude Code sessions. | A promising result on a tiny test set. Rcids wrote both the planted edits and the rules, which may flatter the result; it is not an independent estimate of real-world error rates. |
| Skill picker | 3 of 3 correct on Rcids’s roster: two matches and one correct “none.” | Evidence only for that three-case roster, not for a broad range of prompts or installed skills. |
| Review precheck | A rename-only diff went to fast review. A diff containing a hardcoded key, swallowed exception, and emptied test file went to full review with the right flags. | These examples show intended routing on selected diffs; no larger accuracy rate is reported. |
| Browser navigator | 5 of 5 steps on a synthetic login page. | Rcids says it was not tested against a real browser session. |
| File discovery | Top-three hits were 4 to 6 of 8 across variants, compared with 4 of 8 for plain keyword counting; identical reruns differed by as many as 2. | The author’s small evaluation found no advantage over the keyword baseline and showed run-to-run variation. |
| Latency and input | About 1 second per prompt or edit, about 2 seconds when a violation is confirmed, and about 5,000 input tokens per edit with 20 rules. | These are implementation-specific figures reported by Rcids, not general service guarantees. The model’s individual responses were described as taking tens to hundreds of milliseconds; end-to-end plugin operations took longer. |
| Rule calibration replay | On a different project, 20 rules replayed over 24 real hunks produced no fires above 0.35. | Rcids notes that this fits two opposing explanations: the rules may fit a clean history well, or the rulebook may be blind. No-fires alone cannot distinguish them. |
Rcids also cites a one-week field report on a similar skill router in which about 5% of suggestions were followed by the agent. The year of that report is not stated in the retrieved article, and the figure is not a measurement of jev-tools. Rcids gives it as motivation for leaving the skill hook off by default.
How this approach compares with other ways to make agent decisions
The report does not test competing products or establish that a yes/no model is generally better than an ordinary language-model call or deterministic rules. It does support a narrower comparison:
| Approach | What can be said from this experiment |
|---|---|
| Typed model decision plus code thresholds | The output is structured for code to consume, while thresholds and resulting actions remain explicit in the plugin. The reported rule and browser examples also show that scores can fluctuate around a decision boundary. |
| Ordinary language-model call | No direct comparison was reported. The plugin’s typed responses avoid relying on generated prose for these decisions, but the experiment does not establish an accuracy or speed advantage over ordinary calls. |
| Deterministic keyword rules | The direct comparison was limited to file discovery: top-three hits ranged from 4 to 6 of 8 for the plugin variants, versus 4 of 8 for keyword counting, with reruns varying by as many as 2. That sample does not show a reliable advantage. |
For deployment, the most consequential differences are not just output format. They include repeatability, the cost and latency of additional checks, how each feature behaves during an outage, and what project context leaves the machine.
What leaves the machine
The plugin sends feature-specific context to api.codiv.ai. Depending on the feature, that may include a filename and diff, project rules, a prompt and installed-skill descriptions, excerpts from candidate files, a git diff, or a browser goal, URL, and interactive element names. Shadow mode still sends this data because the plugin needs a response to log what it would do. Only turning the plugin off prevents those API calls.
Rank #4
Rcids says local pattern-based redaction runs before requests and covers common key and credential patterns; files with secret-like names are excluded. The author also warns that pattern matching is not a guarantee. It can miss unusual token formats and does not catch names, email addresses, or customer and employee data; the README also notes that internal business information may not be caught.
Rcids says the public API documentation did not explain storage. The project README describes retention as unknown and advises checking the provider’s terms before sending anything that would not be safe to paste publicly. The available account does not establish that submitted data is not retained, or that redaction makes sensitive projects safe to send.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Setup and a cautious rollout
The documented setup requires Python 3.10 or later on PATH and a free OpenJev key. Hooks invoke python, so a system that exposes only python3 may need an alias. The README recommends storing OPENJEV_API_KEY at the user level rather than committing it to a repository.
JEV_MODE defaults to shadow. The documented defaults are a 0.80 rule-flag threshold and a 0.70 second-look confirmation threshold; these are defaults, not values demonstrated to be optimal. The README describes a free tier of 100 million input tokens, but quota and service terms can change, so check current terms before relying on that allowance.
Best Value
- Start in shadow mode. This lets the hooks log proposed decisions without enforcement, but it does not prevent project context from being sent to the hosted API.
- Review the logs. Look for false positives, missed cases, unstable decisions near thresholds, and behavior on the kinds of edits your project actually makes.
- Calibrate rules before enforcement. Use the rule-calibration feature and inspect its labels rather than treating a quiet replay as proof that the rules are safe or useful.
- Choose activation feature by feature. Consider each feature’s failure policy and the consequences of an incorrect block, injected skill, or review route. Do not infer that success in one feature validates another.
Rcids recommends spending a week in shadow mode, then reviewing logs and calibrating rules before switching to active mode. That is the author’s rollout recommendation, not a tested guarantee that one week is sufficient for every repository.
What the experiment establishes—and what it does not
The useful result is architectural: a model can return a typed decision, while ordinary code owns thresholds and actions. Rcids’s second-look check improved the reported rule-enforcer sample, and the feature-specific failure policies make the consequences of outages explicit. But a few hand-built cases cannot establish real-world accuracy, and fluctuating probabilities can still complicate decisions near a threshold.
The limits matter as much as the promising cases. File discovery did not beat keyword counting in the small evaluation; browser navigation was tested only on a synthetic page; and some features were exercised by script rather than through the actual skill loader. The author’s own summary is apt: “The samples are small and run-to-run noise is real, so read these as smoke tests, not benchmarks.” — Rcids, the jev-tools author, in the DEV Community article.
Rcids describes the project as independent and not affiliated with Codiv, OpenJev, or TypeSafe AI. The account supports conclusions about what the author built, observed, and measured—not independent validation of model calibration, broad performance, provider data retention, or current commercial terms.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

