What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

Neither Jev nor a fine-tuned encoder wins every classification task. In Ikkun’s 2026 experiment on three Japanese datasets, a 310-million-parameter encoder beat Jev on long-document topic classification, while the two systems were statistical ties on two polarity tasks. The practical lesson: test on held-out examples from your own task before choosing a model.

What did the 750-row comparison test?

Ikkun compared zero-shot systems with supervised models using 250 rows apiece from three Japanese classification datasets. The tasks differed in both text length and what the label required:

  • livedoor news: nine-class topic classification; documents averaged 1,174 characters.
  • Rakuten reviews: two-class polarity classification; reviews averaged 138 characters.
  • ChABSA: three-class financial-sentence polarity; sentences averaged 92 characters.

The supervised encoder was sbintuitions/modernbert-ja-310m with a classification head. The zero-shot comparison included Jev 1.13.0. Other tested systems included a local Gemma 4 26B-A4B model in Q4_K_M, SemIf (Qwen3.5-4B), GLiClass multilang-mini, and Laya; a character n-gram plus logistic-regression model served as a simpler supervised baseline. The results below compare Jev and the encoder, not every system in the experiment.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Which system scored higher on each task?

Dataset and task 310M encoder Jev Author-reported paired result
livedoor, nine-class topic 88.8% 76.8% Encoder ahead by 12.0 percentage points; McNemar p=0.00007
Rakuten, two-class review polarity 92.8% 94.4% Jev ahead by 1.6 points; p=0.50, reported as a tie
ChABSA, three-class financial polarity 75.2% 74.0% Encoder ahead by 1.2 points; p=0.83, reported as a tie

These are Ikkun’s measurements, not an independent benchmark. In this sample, the encoder’s clear advantage was on livedoor; the small score differences on Rakuten and ChABSA were not decisive. A nonsignificant result with only 250 examples per task does not establish that the systems are equivalent or that no real difference exists.

How were the results produced?

The encoder used five-fold out-of-fold evaluation. In each fold, it trained on 200 rows and predicted the 50 held-out rows; across five folds, every row received a prediction from a model that had not trained on it. The author says systems were evaluated on identical row IDs and compared with McNemar’s paired test, which evaluates disagreements between paired predictions.

This is a useful way to compare systems on a small fixed sample, but the published account is the experimenter’s report, not an independent audit of the data or implementation. It does not turn three datasets into a broad estimate for other languages, domains, or production traffic.

Why did the winner change with the task?

Ikkun’s interpretation is that topic classification can be strongly signaled by vocabulary: a supervised model may learn that words and phrases predict a category from a few hundred labeled examples. Polarity classification asks for a judgment about meaning, and in these two tested polarity tasks the trained encoder did not show a decisive advantage over Jev.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

That is a plausible explanation of these results, not a general rule that topic classification always favors fine-tuning or that decision APIs always handle sentiment better. The datasets, label definitions, language, and sample all matter.

The simpler baseline is a useful warning

The character n-gram plus logistic-regression baseline scored 88.4% on livedoor, 73.6% on Rakuten, and 64.0% on ChABSA. It nearly matched the encoder on livedoor, but trailed on both polarity tasks. Because the author tested only one configuration, these figures do not establish the best performance a character-based model could achieve.

What should you choose if you have a few hundred labels?

  • No labeled examples, and the task requires a nuanced judgment: A zero-shot service such as Jev is one option to evaluate. The experiment does not establish that it will win on your data.
  • A few hundred labels and a vocabulary-heavy topic task: Try a small supervised encoder and evaluate it on held-out examples from your own domain. The livedoor result supports testing this route, not assuming it will reproduce the same gain.
  • A few hundred labels for polarity: Do not assume fine-tuning will beat a zero-shot system. Jev and the encoder were statistical ties on both polarity datasets in this comparison.
  • Neither approach performs well enough: Inspect label quality and examples of errors; consider sending uncertain cases to a stronger model or a human reviewer. This is the author’s suggested workflow, not a separately measured result.
  • Using Jev-generated labels to train a student: The student can learn the teacher’s errors as well as its useful patterns. If the goal is to outperform the teacher, uncertain examples need better labels rather than simply more labels from the same teacher.

For a fair decision, compare accuracy on your own held-out data alongside latency, local versus hosted operation, privacy and network requirements, labeling and inference costs, confidence calibration, and the consequences of mistakes. The benchmark directly informs only a few of those factors; the rest depend on your deployment.

How much faster was the trained encoder?

On Ikkun’s tested setup, the reported median latency per item was 0.10–0.45 seconds for the encoder and 1.9–2.4 seconds for Jev. These are setup-specific measurements, not service guarantees; they should not be treated as expected latency on other hardware, with other network conditions, or under production load.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The author reports using a Ryzen 9 7940HS mini-PC with Radeon 780M graphics. For the longer-document cross-validation run, training took 112 minutes on CPU and 44.5 minutes on iGPU; for short-sentence training, it took 21 minutes on CPU and 23 minutes on iGPU. Those timings describe one machine and software setup, not hardware requirements or a retail recommendation.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What are the limits of the comparison?

  • Each task had only 250 rows, so the scores may not predict performance on a larger or different sample.
  • Livedoor documents were truncated at 512 tokens, which can omit information from long articles.
  • The encoder used raw argmax confidence, not temperature-scaled or otherwise calibrated confidence. Its confidence values should not be assumed reliable for thresholding without calibration.
  • Only one character n-gram baseline configuration was tested.
  • All three datasets were Japanese and represent particular topic and polarity tasks. The results do not establish performance for other languages, domains, current system versions, or real-world distributions.

Ikkun’s practical summary is: “Measure your task’s shape before you pick a winner — and be suspicious of any ‘open-source Jev replacement’ benchmark where the model was trained on the benchmark.”

Sources and volatile service details

The experimental design, scores, interpretation, and limitations are from Ikkun’s 2026-09-21 report in Agent Journal. It also reports Jev at $0.042 per 1 million input tokens, with output free and an approximately 32k context. Those are the author’s reported service details, not a guarantee of current terms; check the provider’s current pricing and capabilities before budgeting or deployment.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.