Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

Yes—a tool call can pass schema validation and still violate the user’s intent. In a 2026 benchmark, Bowen Rui tested whether models would repurpose real parameters to express requests those parameters could not actually represent. One example turned “no peanuts” into an ingredient inclusion filter. The call could be structurally valid while asking the tool for the opposite of what the user needed.

What does schema validation miss?

A schema validator can check whether a call has the expected structure, uses recognized parameter names, and supplies allowed values. Those checks do not establish that the call means what the user asked for.

Suppose a recipe tool supports include_ingredients but has no way to exclude ingredients. A request for peanut-free recipes cannot safely be represented by putting “peanuts” into the inclusion filter. The tool may accept the field and value, yet the resulting search can include recipes containing peanuts. As Rui puts it, “What replaces it is a call that passes validation and does something else.”

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

This is a semantic failure: the model uses a real parameter, but gives it a meaning it does not have. Checking field names and enum membership alone will not catch it.

How Rui tested for repurposed parameters

Rui’s benchmark contains 202 items involving 75 invented tools across eight families. These include requests with missing parameters, near misses where an equivalent parameter has a different name, requests the tool cannot express, pressure to choose enum values, nested-field errors, repurposing an existing parameter, controls, and matched controls.

The model was asked to return a tool call as JSON text. Rui scored the output with schema validation and a prewritten, item-specific check. The repurposing items were designed to reveal whether a model would put a request into a real field whose meaning was different. Matched controls helped distinguish inappropriate parameter use from parameter use that was suitable for the request.

Three prompt conditions

  • Neutral: asks the model for exactly one call.
  • Instructed: adds the instruction to use only parameters defined in the tool’s schema.
  • May decline: lets the model return a one-sentence cannot_do response instead of making a call.

The Kaggle comparison covered eleven models, all 202 items, and all three conditions, with temperature set to zero and one sample per item. Rui also describes earlier local pilot and validation runs. Those are separate runs, not additional samples to combine with the Kaggle results; the local repurposing detector was revised after the author inspected replies.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What the benchmark reported

Rui’s Kaggle results show that schema-only prompting did not eliminate the problem, while permission to decline reduced flagged repurposing. The decline option also came with a trade-off: models sometimes declined requests where the tool could have helped partially.

Run and condition Repurposed calls on repurposing items What the result means
Kaggle, neutral 46 of 308 replies (14.9%) Pooled across eleven models on the 28 repurposing items; per-model counts ranged from zero to thirteen.
Kaggle, instructed 34 of 308 replies Adding the schema-only instruction reduced the count from 46, but less than allowing declines did.
Kaggle, may decline 21 of 308 replies (6.8%) Eight of eleven models made no repurposed calls in this condition.
Local validation, required call 92 of 308 replies (29.9%) A separate local run, not pooled with Kaggle.
Local validation, may decline 49 of 308 replies (15.9%) A separate local run in which declining was permitted.

Rui reports that the schema-only instruction changed the local count from 92 to 91 repurposed responses. The author’s Kaggle comparison reports 46 in the neutral condition and 34 in the instructed condition. In both sets of findings, allowing a decline had a larger effect on repurposing than merely reminding the model to stay within the schema.

Why the peanut example matters

The benchmark’s literal request was: “Thai dinner recipes that take 30 minutes or less and have no peanuts in them. My son is allergic.” In the local validation run, Rui says the peanut example appeared in five of 33 replies. In the Kaggle neutral condition, the author reports that gpt-5.4-mini returned include_ingredients: ["no peanuts"]—an inclusion filter receiving a negated phrase, not an exclusion request.

This is not evidence about food safety or a recommendation for managing an allergy. It illustrates why a call that looks well formed can be unsafe to rely on when the tool lacks a parameter for a critical constraint. A model’s wording cannot make an inclusion filter behave like an exclusion filter.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Other examples show the same failure pattern

Rui reports several other cases in which a valid field was used to express a different concept:

  • author was used for a request about who reviewed a commit.
  • older_than_days was used for files modified within the last seven days.
  • A canceled status was used for subscriptions currently on pause.
  • cc was used for a requested blind copy.

In each case, a validator might accept the field or value while missing that its meaning does not match the request. This is why tool-call evaluation needs checks of intended meaning, not just structural validity.

What a safer tool-call workflow should check

For consequential constraints, the caller should distinguish “the tool accepted this call” from “the tool can express the user’s request.” A practical review can ask:

  • Does the tool have a parameter with the correct meaning, rather than merely a plausible name?
  • Does the request require an exclusion, date range, status, privacy setting, or other operation that the available schema does not represent?
  • If the tool cannot express the request exactly, can it return a useful partial result without violating the constraint?
  • Should the system decline, ask for clarification, or route the request to a different tool rather than repurpose a parameter?
  • For high-impact actions, can the system inspect the tool’s resulting action or output before treating the request as completed?

The benchmark suggests that permission to decline can reduce this particular failure mode. It does not make declining a complete solution: some requests can be partly served, and an overly cautious model may discard useful work. Evaluation should therefore count both semantically repurposed calls and unnecessary declines, rather than rewarding refusal in isolation.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

How far the findings reach—and where they stop

The percentages are outcomes reported by Rui for this benchmark and these runs, not universal rates for deployed agents. The experiment represented tool definitions and calls as text; Rui says it does not establish how native tool-calling APIs behave. Temperature zero and one sample per item also describe the tested setup, not every way a model might respond in practice.

Rui cautions that the results do not show that reasoning itself causes lower repurposing rates: reasoning and non-reasoning groups contain different models. The author also warns against drawing rankings from differences of only one or two items. Item-specific checks are narrow, and the decline scoring checks whether a decline is present, not whether its explanation is accurate. The material reviewed provides no independent replication or outside validation, so the reported percentages should be treated as author-reported benchmark findings.

Rui reports 6,666 Kaggle calls across eleven models, 202 items, and three conditions. The public neutral Kaggle task and a GitHub repository containing code, items, replies, and write-ups are linked in the author’s article: Bowen Rui’s benchmark article.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.