Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

In Bhushan Kinge’s 2026 benchmark of federal IT solicitations, the hosted typed-decision system Jev narrowly beat Qwen3.5-35B-A3B on primary-class accuracy. More importantly for automation, Jev’s confidence scores supported a cutoff whose Wilson 95% lower bound met a 95% precision target; no reported Qwen or Laya confidence cutoff did. Qwen led a separate fulfillment-mode task. The results suggest that an automation system’s confidence and review policy can matter as much as its raw accuracy—but they do not establish a universal model winner.

What the benchmark tested

Kinge evaluated three approaches on 12,000 U.S. federal IT solicitations, or requests for quotation (RFQs), arriving through SEWP, GSA MAS and GSA 2GIT. The intended workflow classifies an opportunity so a reseller can route it to work such as distributor price lookup, an engineer or original-equipment-manufacturer configurator, publisher authorization, or a statement-of-work process.

Approach Configuration in the benchmark How its output was used
Jev, TypeSafe System One, version 1.13.0 Hosted API using typed questions and calibrated probabilities Jev and Laya received byte-identical typed-question bundles.
Qwen3.5-35B-A3B-FP8 On-premises inference through vLLM It returned a strict JSON schema, which was mapped into the shared taxonomy.
Convai Laya, 421 million parameters Open weights run on an RTX 2000 Ada laptop GPU It received the same typed-question bundle as Jev.

The labels covered purchase type, lifecycle, solution domain, hardware fulfillment mode and flags such as insufficient notice text, RFI, or brand-name-only. The study compared classification and routing signals, not whether a solicitation was ultimately fulfilled correctly.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How to read the accuracy results

The benchmark author reported primary-class accuracy on 741 shared rows with one unambiguous gold class. These are Kinge’s benchmark results, not independent validation or an industry-wide estimate.

System Primary-class accuracy Evaluation subset
Jev 91.9% 741 shared rows with one unambiguous gold class
Qwen3.5-35B-A3B 89.6% Same 741-row subset
Convai Laya 78.0% Same 741-row subset

That subset is much narrower than the 12,000 solicitations initially sampled. Of those, 927 had a quote, and 741 had a single unambiguous primary class. The composition labels came from product types on the latest quote lines when the reseller’s sales team quoted an opportunity. They therefore reflect downstream sales handling, not objective ground truth for every solicitation, and favor opportunities that were pursued and quoted.

The gold subset was also heavily weighted toward Hardware: it made up 77% of single-class rows. Services had six rows and Maintenance & Support had 32, so results for those small classes are especially uncertain. The study did not have human gold labels for subclass, lifecycle and solution-domain decisions.

Why confidence changed the automation picture

Accuracy answers how often a system is right on a labeled set. For automation, another question matters: can the system identify a subset of its answers reliable enough to accept without human review? Kinge reported Jev’s expected calibration error (ECE) as 0.049 and evaluated precision-versus-coverage cutoffs using the Wilson 95% lower bound.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Confidence approach Reported coverage and precision What it implies for a 95% precision target
Jev, cutoff 0.94 Accepted 641 of 741 rows (86.5% coverage) at 96.7% observed precision The Wilson 95% lower bound stayed at or above the 95% target along the reported cutoff envelope.
Qwen, prompt-defined “high” category Covered 97.8% of rows at 90.1% precision No Qwen confidence bucket met the 95% target.
Laya confidence values No useful target-meeting cutoff was reported. The confidence values did not provide a usable threshold for the target.

Jev’s result is not simply “96.7% is better than 90.1%.” The operational distinction is that, in this labeled subset, Jev could mark a large share of rows for acceptance while satisfying the study’s lower-bound criterion. That makes its confidence score useful as a proposed review boundary in this experiment. It does not guarantee equivalent precision on a different reseller’s work, a later period, or an unreviewed class distribution.

Qwen led the separate fulfillment-mode task

On 634 rows labeled for hardware fulfillment mode, Qwen scored 71.0% accuracy, Jev 65.0%, and Laya 45.7%, according to Kinge’s benchmark. This ranking should be treated cautiously: fulfillment labels were inferred from configurator fingerprints and distributor information on quote lines, and the author described the “eight or more lines from one OEM” rule as an unvalidated heuristic requiring human validation.

Within the same task, configured-build precision ranged from 19% to 36%, and the “mixed” category was effectively unsolved. A stronger overall fulfillment-mode accuracy therefore should not be mistaken for reliable performance on every fulfillment category.

Queue policy can outweigh a flag’s apparent caution

The systems also emitted flags that could be used to route cases for review. In Kinge’s queue simulation, sending every flagged row to human review automated 26.6% of volume at 93.5% precision. Treating flags as attributes rather than automatic blockers automated 91.9% at 93.8% precision on the simulation’s scored rows. The two precision figures came from different scored row counts, so the comparison is suggestive rather than a controlled, same-denominator comparison.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The brand-name-only flag fired on 54% of Jev rows. Making every such flag a hard stop sharply reduced simulated automation without an observed precision gain in the reported simulation. That does not mean flags should be ignored: it means a flag’s operational meaning should be designed and validated rather than assumed. A flag can inform prioritization or add context for a reviewer without automatically blocking a case.

Agreement, variants and information limits

Jev and Qwen agreed on 91.0% of 11,931 paired outputs. Kinge reported 1,068 disagreements, concentrated around category boundaries such as Hardware versus Other and Hardware versus Software. Those disagreements are useful candidates for stratified blind human adjudication: reviewing only disagreements can expose boundary problems, while sampling across categories can help detect errors shared by both systems.

Nine Jev variants scored between 91.2% and 91.9% on the same 741 rows, with 676 to 681 correct. A tested excerpt of up to 1,500 attachment characters was available on only 5.6% of rows and did not measurably improve this task. That finding applies to this limited excerpt test, not to richer attachment extraction or broader document understanding.

For the 12,000 inputs, Kinge’s report lists Jev at $0.78 using list input-token pricing, with 185 ms p50 latency, 273 ms p95 latency and zero errors. Qwen used roughly 80 GPU-minutes on a shared cluster, processed 2.7 rows per second and returned 69 permanently malformed responses. Laya ran locally on an RTX 2000 Ada laptop GPU, with 299 ms p50 and 576 ms p95 latency and zero errors. These are results from different products, hardware and deployment assumptions; they are not a like-for-like price comparison.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What the sample and method do—and do not—establish

  • The sample comprised 6,000 SEWP, 3,000 GSA MAS and 3,000 GSA 2GIT records created from November 8, 2024 through September 22, 2026. The order was deterministic by md5(id). This construction does not make the sample representative of all federal procurement.
  • Qwen produced 69 permanently malformed responses among the 12,000 inputs. Those rows were excluded from paired metrics, leaving 11,931 paired rows; because the malformed cases were not random, excluding them may affect comparisons.
  • Qwen’s three confidence levels were imposed by the prompt, while Laya used shipped defaults, one checkpoint, single-row execution and no threshold tuning. Confidence comparisons therefore reflect these configurations, not every possible deployment of either model.
  • The benchmark covers one organization, one domain, one Jev version and one measurement period. Public aggregate results do not include the underlying solicitation sample, quote identifiers, gold files or per-row predictions, limiting independent reproduction.
  • Kinge identified blind review of a stratified sample and 300 Jev–Qwen disagreements, human validation of fulfillment labels, and drift regression as follow-up work. These were pending or planned, not completed evidence.

What an ML team can take from the results

The benchmark supports a practical decision framework, not a universal ranking. If a team is considering automation, it should measure the dimensions that determine how errors reach people and workflows:

  • Match the evaluation to the decision. Separate primary classification from fulfillment mode and other routing flags. Do not treat success on one task as evidence for another.
  • Audit the labels. Verify that quote-derived labels reflect the decision the automation is meant to make, especially for small classes and fulfillment categories.
  • Choose a review boundary from held-out data. A confidence cutoff is valuable only if confidence ranks reliability well enough, and its precision remains adequate on representative, independently checked cases.
  • Define flags as workflow signals. Decide which flags trigger a hard stop, which change priority, and which are merely recorded as attributes; then evaluate the resulting queue, not just the classifier.
  • Track operational failures as model outcomes. Malformed output, latency, throughput and deployment constraints can determine how many rows actually complete without human repair.
  • Review disagreement and drift. Blindly adjudicating sampled disagreements can reveal taxonomy boundaries; repeat evaluation as solicitations and downstream processes change.

The central lesson is that confidence can change which system is useful for a particular automation policy even when the raw-accuracy gap is small. Kinge’s results make Jev the stronger candidate for primary-class automation under the tested confidence criterion, and Qwen the stronger candidate on the tested fulfillment-mode labels. The label limitations, unequal configurations and single-reseller scope prevent either finding from being generalized into a universal winner.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.