Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Clear out junk files and repair common Windows errorsFree Scan →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
In Zika Zag’s 53-task case study, Claude Opus 5 made more correct project assignments than Jev 1.13 (33 versus 30), while Jev made fewer wrong assignments (9 versus 12) and left more tasks unassigned (14 versus 8). The author measured Jev’s median response time at 209 ms, compared with 4,689 ms for Opus. Those results describe one person’s tasks and one implementation—not a general model ranking.
What the 53-task comparison found
Zika Zag tested project assignment for 53 personally selected tasks whose correct projects were known. The figures below are the author’s reported results, not independently verified measurements. “Right when answered” excludes blank responses; a blank counted as neither right nor wrong.
| Model | Right | Wrong | Blank | Right when answered | Median latency | Cost per call |
|---|---|---|---|---|---|---|
| Jev 1.13 | 30 | 9 | 14 | 77% | 209 ms | About $0.00019 median, measured from 53 calls |
| Claude Opus 5 | 33 | 12 | 8 | 73% | 4,689 ms | About $0.025, estimated from five calls |
| Claude Sonnet 4.6 | 29 | 15 | 9 | 66% | 1,758 ms | Not measured |
| Claude Haiku 4.5 | 19 | 8 | 26 | 70% | 1,013 ms | Not measured |
For this task set, Opus produced three more correct assignments than Jev. Jev had three fewer wrong assignments, but it also abstained on six more tasks. The “right when answered” percentage should be read alongside those counts: a system can improve that rate by answering fewer tasks.
PC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteWhy a blank can be better than a wrong project
In Open Walnut, a wrong assignment can file a task under a project the user is not currently viewing. A blank leaves it in the Inbox, where it remains visible for manual sorting. That makes the error trade-off specific to the app’s workflow: fewer wrong assignments may matter more than maximizing automatic filing, depending on how costly it is to review an Inbox item.
#1 Best Overall
The comparison treated three notes that seemed equally compatible with two projects as wrong for every engine. This is a scoring choice, not proof that one project was objectively the only defensible destination. It also means the numbers reflect the author’s project labels and rules for judging ambiguous notes.
How the test was run
Same production path, different model routes
Jev used Open Walnut’s production parseQuickTask code through OpenRouter. The Claude models received the production quick-parse prompt and the same project digest through Bedrock. The reported table covers project assignments only; it does not present a separate scored comparison of priority or tier suggestions.
Rank #2
Project context and answer choices
The digest included the 20 projects with the most open tasks, while the choice question offered every project. Zag reports that adding each project’s short summary—up to 160 characters—as evidence for its option helped represent projects outside the top-20 digest. That is an implementation detail from this run, not a controlled test of how much the summaries improved accuracy.
Do these 3 things before closing this tab:
1Scan for outdated or missing drivers - takes under a minute2Repair Windows errors before they cause bigger problems3Fix the driver behind crashes, sound loss and screen glitchesConfidence thresholds and prompt wording
Jev returned a selected option, option probabilities, and confidence rather than free-form text in this workflow. The project confidence floor was 0.4; tier and priority floors were 0.5, and moving a session used a 0.6 floor. When confidence fell below a field’s floor, the field stayed blank.
Rank #3
The author reports that lowering the project floor from 0.5 to 0.4 added nine right assignments and one wrong assignment on this set. In a production-path comparison with revised wording, the reported counts changed from 22 right, 8 wrong, and 23 blank to 30 right, 9 wrong, and 14 blank. Because both the wording and floor are described as changing, these figures do not isolate the effect of either change on its own.
Zag also says the earlier question wording emphasized “none” and left generic bug reports blank. The revised version framed bug reports, feature ideas, and investigations as usually belonging to a product project whose scope covers them, with “none” reserved for personal errands and reminders. This is the author’s observation about this implementation, not a general finding about task-classification prompts.
Rank #4
Speed and cost: useful figures with different caveats
The author measured Jev’s median latency at 209 ms and Opus’s at 4,689 ms. These are results from the described setup, not a guarantee of response times for other networks, providers, prompts, or dates. In a separate run of 92 unlabelled to-dos, all Jev calls returned; median latency was 162 ms and the 95th percentile was 489 ms. Since those tasks had no known answers, that run speaks to call completion and timing, not assignment accuracy.
The reported Jev cost—about $0.00019 per call—is the median usage.cost across 53 calls. Opus’s roughly $0.025 figure is an estimate based on five calls, approximately 4,290 input tokens and 130 output tokens, cited rates of $5 per million input tokens and $25 per million output tokens, and no prompt caching. The two figures were calculated differently, so they are not an audited like-for-like cost comparison. Sonnet and Haiku costs were not measured.
Best Value
Abstention, failures, and fallback behavior
The implementation distinguishes an intentional blank from a failed request. Low confidence or a “nothing fits” response leaves the field blank. A network error, malformed response, or missing confidence triggers a fallback to the prior fast-model path. The quick-add call also has a 2.5-second limit. These safeguards describe Open Walnut’s reported behavior; they do not establish how another application using the same model would handle failures.
What this case study can—and cannot—show
The results answer a narrow question: how these configured systems performed on Zag’s 53 known-project tasks using this production path and scoring approach. They do not establish which model is best for task classification in general. A different project list, writing style, ambiguity policy, prompt, confidence floor, or provider route could change the balance between correct assignments, mistakes, and abstentions.
For developers considering this pattern, the practical lesson is to measure the three outcomes separately: correct assignments, wrong assignments, and blanks. Latency is valuable only alongside that error profile, and confidence thresholds should be chosen according to what happens to an item when the system declines to decide.
Setup details reported by the author
Zag’s article names the model identifier typesafe/jev-1.13, the endpoint https://openrouter.ai/api/alpha/decisions, and the Open Walnut path Settings → Tasks → Smart task creation, with Jev selected under “Uses.” The article also describes a jev: configuration section. These are date-sensitive implementation details reported on September 24, 2026; current interface labels, endpoint availability, model availability, and pricing are not independently established here.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

