Compare small language models (SLMs) on the same held-out examples, with the same instructions, schema, decoding settings, and production output mode. Measure whether each model makes the right decision separately from whether its response parses or passes schema validation. For tool use, also score tool choice, arguments, and successful execution. There is no evidence here for a universal SLM winner: the best choice depends on your task and deployment.
Define what a correct decision means
Before running models, turn the task into something an evaluator can judge. Specify the input, allowed decisions, required output fields, and the conditions for success. For classification, list the valid labels; for extraction, define which values count as correct; for routing, specify the correct destination and when no route is appropriate.
Make abstention and clarification explicit. A tool-oriented task may require the model to call a tool, decline to call one, ask for missing information, or select a different tool. Treat these as distinct expected outcomes rather than leaving the model or evaluator to infer them.
OpenAI’s evaluation guidance recommends checking instruction following, functional correctness, tool selection, data precision, and agent handoff where relevant. The precise criteria should reflect the application, not a generic idea of a good answer.
The Tool Desk
Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →#1 Best Overall
Build a representative held-out test set
Use examples that reflect the inputs the application will actually receive. Include routine cases, ambiguous or incomplete inputs, and consequential edge cases. Run the same cases against every candidate. Keep a held-out set for final comparison so prompt or schema changes are not judged only on examples used to tune them.
There is no universally adequate sample size established by the cited guidance. Choose a set that is large and varied enough for the intended workload, and report its size and limits alongside the results. A score on a narrow or unrepresentative test set should not be presented as a reliable estimate for a broader deployment.
Hold evaluation conditions constant
For a fair comparison, fix the task instructions, schema, available tools, decoding settings, and retry policy. Record the model and configuration, too. If your application will use a provider’s constrained-output feature, evaluate that feature in the candidate setup; if you are choosing between it and prompt-only JSON, test both modes rather than attributing every difference to the model’s weights.
Output mode matters. OpenAI’s documentation distinguishes JSON mode, which ensures valid JSON, from Structured Outputs, which are designed to ensure adherence to supported schemas and models. Function calling is intended to connect a model to tools or APIs; a structured response format shapes the model’s answer. These modes serve different purposes, and the mode used in production belongs in the evaluation.
Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallCrashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteScore decision quality separately from formatting
A response can pass a parser or schema validator while encoding the wrong choice. Track these outcomes independently rather than treating successful formatting as task success.
| Measure | What to check |
|---|---|
| Decision accuracy | Whether the selected label, route, value, or action matches the expected decision. |
| JSON parsing | Whether the response can be parsed as JSON. Parsing alone does not establish that it follows the required schema. |
| Schema validity | Whether the parsed response satisfies the target schema. |
| Semantic validity | Whether the field values are correct and consistent with one another, even when the response passes schema validation. |
| Wrong-but-valid rate | How often outputs satisfy the schema but encode an incorrect decision. |
For objectively checkable tasks, use exact-match comparisons or executable checks where possible. For judgments based on multiple criteria, write those criteria down and apply them consistently. Do not collapse decision accuracy, schema validity, and semantic validity into a single pass rate.
Rank #3
Evaluate tool use through execution
For tool-oriented tasks, score more than whether a call is well-formed. Check whether the model chose the right tool, supplied precise arguments, and handed off or declined appropriately. Where feasible, run calls in a safe test environment and measure whether the intended task completed.
A 2026 paper by Jaideep Ray, The Constraint Tax, illustrates why this separation matters. In its deterministic calendar tool-call task using Qwen2.5-1.5B, prompt-only JSON and the tested hard tool-call schema both had 100.0% schema validity, but executable accuracy was 91.5% and 48.0%, respectively. Those are results for that model, task, and comparison—not expected rates for other applications.
Check repeatability and operating fit
Generative systems can return different outputs for the same input. OpenAI’s evaluation documentation warns that this variability makes traditional software testing alone insufficient. Repeat runs when that variability could change a decision, and include varied cases rather than relying on a single successful output.
Measure latency and cost under representative conditions if they affect deployment. Treat them as application-level measurements: the cited sources establish no universal acceptable threshold. A faster or cheaper model may be a poor fit if its errors create extra review or failed actions; a higher aggregate benchmark score does not by itself establish suitability for a narrow task.
Use benchmarks as supporting evidence
Public benchmarks can help characterize constrained output and tool use, but they do not replace tests on the application’s own inputs and decisions.
| Benchmark | What it helps assess | What its result does not establish |
|---|---|---|
| JSONSchemaBench | Constrained-decoding efficiency, coverage of constraint types, and output quality. Its 2025 paper describes 10,000 real-world JSON schemas paired with the official JSON Schema Test Suite. | Whether a model will make the correct semantic decision on your workload. |
| BFCL V4 | Function-calling and agent behavior, including multiturn interactions. Stanford HAI’s 2026 AI Index reports that agentic tasks account for 40% of its overall score and multiturn interactions for 30%; the remainder is split across live, nonlive, and hallucination categories. | How an SLM will perform on a specific organizational task. The same report describes about a 21-percentage-point spread in overall accuracy among the top 15 models as of early 2026; that figure concerns the reported leaderboard and version. |
Scores from different benchmarks or evaluation setups are not directly comparable without checking their versions, tasks, scoring rules, and configurations.
Best Value
Interpret published results in context
The Constraint Tax reports 15,000 commodity-GPU generations across Qwen2.5-0.5B, Qwen2.5-1.5B, and SmolLM2-1.7B. In its tested hard answer-only schema decoding setup, it reports schema validity ranging from 61.5% to 100.0%, answer accuracy from 19.7% to 11.0%, and wrong-valid-schema outputs from 49.5% to 88.9%. These are findings for the paper’s tested models and setup, not forecasts for other tasks. The paper’s broader lesson is to report schema validity, answer accuracy, executable accuracy, and wrong-valid-schema rate separately.
Choose based on the workload, not one headline score
- Set minimum requirements for decision correctness and reliability based on the consequences of errors.
- Compare candidates on the same held-out cases and production output path, using the same scoring rules.
- For tools, include selection, argument accuracy, and execution success; for structured responses, include schema validity and wrong-but-valid outputs.
- Include repeat-run behavior, latency, and operational cost when they matter to deployment.
- Report the test set, output mode, schema, decoding configuration, number of runs, and scoring rules with the results.
Select a candidate that meets the application’s correctness and reliability requirements under its real deployment constraints. The available evidence supports a repeatable, workload-specific evaluation—not a universal ranking or a hardware recommendation.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

