iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
In Christian Anderson’s 62-case test of claims in product descriptions and posts, adding Jev as a second checker reduced the number of unsupported claims that passed under his chosen combined rule. The rule got 61 cases right, with one false positive and no false negatives in that sample. That is a useful result for his publishing workflow—not proof that two checkers will improve every dataset or domain.
What Anderson tested
Anderson used claims from actual Gumroad product files and his DEV posts, asking whether each claim was supported by its source before publication. The test set contained 62 cases: 22 claims judged true of their source and 40 that went beyond what the source supported. These labels followed from how the cases were constructed; the post does not establish an independently audited ground truth or an independent reproduction from public raw cases and code.
He ran two checkers separately on the cases:
- DeepSeek chat: deepseek-v4-flash, prompted to read the source and answer PASS or FAIL.
- Jev: typesafe/jev-1.13, which returned a probability that a claim was supported.
Jev’s probability needed a cutoff to become a decision: at a 0.5 threshold, a score of at least 0.5 passed; at 0.9, only scores of at least 0.9 passed. Anderson then compared those individual rules with a combined rule.
Free tools Windows power users keep installed
One-click scans. No signup required.
How the checkers performed on the 62 cases
In the table, a true positive (TP) is a supported claim passed, a false positive (FP) is an unsupported claim passed, a true negative (TN) is an unsupported claim rejected, and a false negative (FN) is a supported claim rejected. “No answer” means DeepSeek did not return a decision. The figures are Anderson’s reported results for this test set.
#1 Best Overall
- Covers 10+ AI prompt frameworks (AIDA, PAS, SWOT, SMART Goals, etc.) Easily turn your workspace into the Empire of AI with the AI Prompting Desk Mat, crafted for thinkers, creators, and professionals working with ChatGPT, Copilot, and other AI tools. Made of 3mm thick neoprene material with an anti-slip backing and hemmed edges, this mat offers comfort, durability, and a clean surface for your keyboard and mouse.
- Includes do’s, don’ts, and real-world prompt examples, this isn’t just a desk accessory — it’s a visual guide to mastering AI prompts. Whether you use chatgpt, PromptPerfect, AIPRM, FlowGPT, PromptHero, or any other platform, this mat helps you write effective prompts with proven frameworks and structured thinking. Ideal for anyone learning AI engineering, exploring AI for business, or taking AI training courses, it bridges creativity and precision in every prompt you write.
- Inspired by the best concepts from AI books & ChatGPT guides, it’s perfect for professionals, educators teaching with AI, or beginners curious about how to use AI productively. Boost your skills, enhance your workflow, and create smarter ideas — right from your desk.
- Hemmed sewn edges for a premium, long-lasting finish, paired with Smooth neoprene surface, 3mm thick for comfort and durability
- Size: 12 x 22 inches — fits perfectly under laptop or keyboard
| Checker or rule | TP | FP | TN | FN | No answer | Accuracy (answered) | Mean time |
|---|---|---|---|---|---|---|---|
| DeepSeek chat | 22 | 2 | 35 | 0 | 3 | 96.6% | 21.5 s |
| Jev, pass at p ≥ 0.5 | 22 | 4 | 36 | 0 | 0 | 93.5% | 0.35 s |
| Jev, pass at p ≥ 0.9 | 21 | 0 | 40 | 1 | 0 | 98.4% | 0.35 s |
| Both combined | 22 | 1 | 39 | 0 | 0 | 98.4% | not stated (Anderson, 2026) |
The combined rule’s 98.4% answered accuracy corresponds to 61 of 62 cases right. Its one remaining error was an unsupported claim that DeepSeek passed while Jev scored 0.63. At the 0.9 Jev-only cutoff, there were no false positives, but one supported claim was rejected; the combined rule handled that trade-off differently.
Why Anderson combined the checkers
Their mistakes did not fully overlap. For example, DeepSeek passed a claim that a holiday pricing guide would help readers “save at least £25,” while Jev assigned it a support probability of 0.13. In the other direction, Jev passed four unsupported claims at the 0.5 cutoff. Three were product descriptions that overstated coverage, and DeepSeek rejected those three. That error pattern—not merely the number of checkers—made a fail-if-either-fails policy useful to Anderson.
Rank #2
Anderson’s live policy is to fail a claim if either checker says FAIL. If DeepSeek returns no answer, Jev must score at least 0.8 for the claim to pass. That no-answer fallback is stricter than simply treating Jev’s ordinary 0.5 threshold as sufficient.
The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Speed, repeatability and cost in this run
Anderson reports that he ran all 62 cases through Jev twice. Scores changed by no more than 0.04 and by 0.007 on average. In the same reported run, Jev’s median time was 0.31 seconds compared with 20.6 seconds for DeepSeek; the table’s mean times were 0.35 seconds for Jev and 21.5 seconds for DeepSeek. He reports a total cost of $0.0018 for all 62 Jev checks. These are measurements and costs from that particular run, not present-day pricing or guaranteed latency.
Rank #3
- 𝐑𝐄𝐒𝐄𝐓 𝐘𝐎𝐔𝐑 𝐌𝐈𝐍𝐃 𝐈𝐍 𝟔𝟎 𝐒𝐄𝐂𝐎𝐍𝐃𝐒 – A simple, screen-free way to disconnect after a high-demand workday or regain focus during a busy afternoon. Pull one of these mindfulness cards, pause, and follow a practical prompt designed to bring calm, clarity, and grounding in about a minute—no app, journal, or meditation experience needed.
- 𝐅𝐈𝐍𝐃 𝐓𝐇𝐄 𝐂𝐀𝐋𝐌 𝐘𝐎𝐔 𝐍𝐄𝐄𝐃 𝐓𝐎𝐃𝐀𝐘 – Includes 52 color-coded prompts across Focus, Calm, Gratitude, Self-Compassion, and Presence. These mindfulness cards for adults make it easy to choose the category that fits the moment, or pull a card at random for a quick daily ritual inspired by approachable mindfulness and grounding practices.
- 𝐁𝐔𝐈𝐋𝐃 𝐀 𝐒𝐄𝐀𝐌𝐋𝐄𝐒𝐒 𝐂𝐀𝐋𝐌𝐈𝐍𝐆 𝐇𝐀𝐁𝐈𝐓 – Keep these self care cards on your desk to break the midday work loop, in your bag for travel, or on your nightstand to transition peacefully into sleep. These bite-sized practices fit naturally into work breaks, quiet mornings, evening wind-downs, and everyday wellness routines.
- 𝐌𝐀𝐃𝐄 𝐓𝐎 𝐅𝐄𝐄𝐋 𝐏𝐑𝐄𝐌𝐈𝐔𝐌, 𝐔𝐒𝐄𝐃 𝐃𝐀𝐈𝐋𝐘 – Crafted from thick 350 GSM cardstock with a smooth premium finish, these cards feel substantial in hand and are designed to withstand repeated shuffling, daily handling, and carrying in a bag or desk drawer without easily bending or creasing. Compact 2.5" x 3.5" size makes them easy to keep close wherever life takes you.
- 𝐆𝐈𝐕𝐄 𝐀 𝐆𝐈𝐅𝐓 𝐓𝐇𝐄𝐘'𝐋𝐋 𝐀𝐂𝐓𝐔𝐀𝐋𝐋𝐘 𝐔𝐒𝐄 – Beautifully designed and easy to use, Mindful Reset makes a meaningful gift for mindfulness, meditation, and daily affirmations. Whether used as meditation cards, affirmation cards, or a simple wellness ritual, this thoughtful deck is perfect for women and men, friends, coworkers, teachers, therapists, students, and loved ones looking to bring more calm and intention into everyday life.
What this result does—and does not—say about routing
This experiment is about checking whether claims are supported by source material. It is not a test of Jev choosing the best model for arbitrary prompts. Jev’s routing documentation describes typed outputs such as a finite model choice, complexity score and probability of needing tools. It also describes jev-router as an open-source, OpenAI-compatible LiteLLM proxy that summarizes incoming messages, filters candidate models by capability, then lets Jev choose; the documentation says a rules-based cheapest-eligible fallback is used when no key is set. Jev routing documentation
A separate paired and self-audited evaluation by Jiawei Li, dated October 1, 2026, examined Jev and Laya at 11 agent decision points. Its abstract reports Jev was significantly more accurate on nine points, but neither system beat chance on zero-shot model routing, and both tied on RAG relevance gating. It also says errors in an earlier analysis distorted deployment claims. This is a different benchmark from Anderson’s claim-support test, and it is a reason not to generalize his result to unrelated routing tasks. Li’s October 2026 evaluation
Rank #4
- GO BEYOND SMALL TALK — 52 cards with 104 open-ended questions (two per card) that turn dinners, road trips, and quiet nights in into conversations you'll actually remember. The original Holstee reflection deck.
- TOGETHER OR ON YOUR OWN — spark deeper conversations with couples, families, friends, and coworkers, or use the deck solo as journaling and self-reflection prompts. No rules, no setup — just draw a card and go deeper.
- COLOR-CODED BY THEME — questions span Gratitude, Wellness, Intention, and more, so you can steer toward what matters most in the moment. Inspired by mindfulness and positive psychology.
- SMALL ENOUGH TO POCKET, BEAUTIFUL ENOUGH TO DISPLAY — each card carries a unique, abstract design. Take the deck on the go, or leave it out on the coffee table.
- QUALITY YOU CAN FEEL — made in the USA from sustainably-forested paper with vegetable-based inks and a starch-based laminate that keeps them durable. As kind to the planet as they are to your conversations.
When a second checker may be worth the extra step
Anderson’s result supports testing a second checker when the cost of publishing unsupported claims matters and different tools may catch different errors. It does not establish that adding a checker will improve accuracy on your material. Before relying on such a policy, evaluate it on examples representative of your own publishing workflow and pay attention to:
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Quick Recap
Best Value
- False positives: unsupported claims that still pass.
- False negatives: supported claims that are rejected, which can create unnecessary review work.
- Non-answers: define whether another checker can decide the case and what threshold it must meet.
- Threshold choice: changing Jev’s cutoff altered the balance between missed overclaims and rejected supported claims in Anderson’s sample.
- Repeatability, latency and cost: measure them in the same conditions and for the tools and versions you plan to use.
- Representativeness: a small, specific set of product descriptions and posts may not reflect another subject, writing style or source type.
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

