Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

One direct Jev judgment is not automatically better—or worse—than extracting a dozen scores. In three classification settings, dimension scores helped when a direct answer was weak, but they also produced a much higher false-positive rate on difficult benign security text. A free character-bigram baseline matched the dimension model on the Japanese task, while a combination of methods led on a contextual bookkeeping test. The practical answer is to start with a cheap baseline, test the direct call, and add dimensions only when they improve the errors that matter on representative data.

What the comparison tested

In experiments published by ikkun on 17 September 2026, the direct method asked Jev for one classification per row. The decomposed method asked 12–14 narrower questions, cached the resulting scores, then fitted weights locally using the author’s labeled examples. The comparison also included local baselines such as character n-grams, word-bigram naive Bayes, and a majority classifier. Results came from out-of-fold evaluation, with statistical tests reported for selected differences. Read the experiment and its task details.

The three settings were a synthetic B2B reply task, a Japanese natural-language-inference task in which no single sentence determined the label, and an export from the author’s bookkeeping data. They are distinct tasks, not a controlled test of one universal classification problem. In particular, the bookkeeping setup used contextual fields and the security-text test measured a specific hard-benign subset.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How the methods performed

Synthetic B2B replies: random splits overstated generalization

On the synthetic reply task, the 12-dimension pipeline reached 98.0% accuracy with random row folds, but 90.0% when entire template families were held out. Character-bigram naive Bayes scored 93.5% on those grouped folds. The gap illustrates why a random split can mislead when near-duplicate templates appear across training and test data. The author also judged the task too easy for the direct question to provide strong evidence about model capability.

Japanese NLI: dimensions helped, but a free baseline tied

On the Japanese task, the direct Jev call scored 64.7% accuracy and the 12-dimension model with fitted weights scored 74.0% out of fold. The reported difference was +9.0 percentage points, with a 95% confidence interval of +3 to +15 points and p=0.0050. Character-bigram naive Bayes also scored 74.0%, without API calls. Thus, decomposition improved on the direct call in this experiment, but it did not beat the best reported local baseline.

Interpret the result in context: the labels contained noise and the dataset was skewed toward neutral examples, which made up 55% of the labels. The experiment demonstrates a useful candidate approach for this dataset, not a general advantage for dimension scoring.

Bookkeeping: context and complementary errors mattered

On 2,101 test rows from the author’s accounting data, a stack combining 14 dimensions with 12 n-gram class probabilities achieved 0.9695 accuracy. Word-bigram naive Bayes alone scored 0.9491; dimensions with fitted weights scored 0.9105; and a direct 12-choice call scored 0.3998.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

This was an “import with context” setup: the model used descriptions, amounts, and credit-side account context, and the task focused on the top 12 debit accounts. It was not a bare bank-statement benchmark. The stack’s result is useful because the two component approaches corrected different rows; it does not show that stacking will help when methods make the same mistakes.

Accuracy is not the only deployment metric

Hard benign security text: dimensions raised false alarms

Among 339 benign rows that mentioned injection techniques, the direct choice method had a 1.5% false-positive rate; the 12-dimension model’s rate was 37.2%, about 25 times higher in this test. The author linked the dimension model’s failure in part to surface cues such as obfuscation and hidden content, which also occur in benign security documentation. These measured rates describe this dataset and evaluation, not the expected rate in another deployment.

For a security guardrail, a strong aggregate accuracy score can conceal an unacceptable burden of false alarms. Evaluate representative benign cases that resemble the risky language your system will encounter, and choose thresholds against the cost of each error.

A high confidence score did not ensure a correct answer

On the Japanese task, 126 of 300 rows received direct-call confidence of at least 0.9; that subset was 72.2% accurate. In this experiment, high reported confidence was not a guarantee of correctness, so it should not replace validation against labeled examples.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What the results say about cost

For the reported experiment, ikkun counted 34.1 million input tokens across 5,477 test rows and 25,174 Jev calls, with $1.43 in total input cost at the then-stated rate of $0.042 per million input tokens. Output tokens were counted separately and were not priced in that total. The author estimated $19–26 per million rows for one direct question and $42 for twelve dimensions. Those are the author’s reported workload estimates and pricing assumptions, not independently verified current rates or a guarantee of future API cost.

More dimensions mean more questions and therefore more token use in this setup. That extra cost is worthwhile only if the improvement on your target task and error profile is worth it. A no-API baseline can establish whether paid model calls are necessary at all.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

A practical way to choose between one call and dimensions

  1. Define the decision and its costs. Specify what counts as a false positive and false negative, and choose the metric that reflects the consequences. For a guardrail, include a hard-benign test set rather than relying on overall accuracy alone.
  2. Build a local baseline first. Try a simple majority classifier and an n-gram or TF-IDF approach. Keep train and test data separated in a way that reflects deployment—for example, hold out whole template families when templates may recur.
  3. Evaluate one direct question. Keep the direct Jev call if it meets the required performance and error thresholds. Do not add a decomposition step just because it is possible.
  4. Try dimensions when the direct call falls short. Write narrow questions that capture meaningful cues, fit weights only on training data, and compare results on held-out examples. The experiments used author-written dimensions, so their results do not establish that arbitrary dimensions will work as well.
  5. Compare row-level errors before combining methods. A stack may help if the methods recover different mistakes; it adds little if they fail on the same examples. Test the combined system on the same leakage-resistant split and measure the metric that matters.
  6. Tune thresholds on representative data. Do not assume a default threshold such as 0.5 is suitable. A separate 2026 Jev benchmark reports that binary probabilities could rank examples well even when a fixed 0.5 threshold performed poorly on some tasks. Its findings support task-specific threshold evaluation, not any particular threshold for these three experiments.

How to interpret broader Jev benchmark results

A 2026 preprint by Tobias Deußer, Lorenz Sparrenberg, and Rafet Sifa describes Jev’s typed choice, score, and yes/no question interfaces and evaluates pinned version jev-1.13.0 on 37 datasets and 346,009 requests. The authors note that results may depend on template wording, that each request was run once, and that the vendor’s jev-latest alias can move to newer versions. Read the benchmark preprint.

That benchmark is additional context, not a replication of ikkun’s three experiments: it used a different protocol and a pinned version, while the primary experiment used jev-latest. Neither set of numbers should be treated as a direct validation of the other. The broader benchmark’s one-template-per-dataset and one-run design also leaves prompt and run-to-run variation unmeasured.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.