Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

The available evidence does not establish what happened in the claimed 36 runs. It does not identify the four schema changes, the three frameworks and their versions, the model, or the run procedure. Without those details, there is no sound basis for saying which framework broke, how often it failed, or which one performed best.

What can be compared is narrower but useful: framework documentation shows different surfaces for describing and validating tool inputs, while agent-debugging research treats malformed or schema-invalid calls as a distinct failure type. Those documented behaviors help explain what a schema-change test should measure; they do not verify the experiment’s results.

What can be established about the 36-run claim?

The headline describes a first-person experiment, but the material available for this article does not include its source post, code, run data, or methodology. The number 36 should therefore be treated as a title-supplied claim, not as a verified study statistic. No outcomes or framework ranking can responsibly be reconstructed from it.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

To assess a comparison of this kind, readers would need at least:

  • The exact four schema variants, including which fields, types, requiredness rules, or allowed values changed.
  • The names and versions of all three agent frameworks, plus the model and its relevant configuration.
  • The task prompt, tool implementations, run allocation, and whether each run began from the same conditions.
  • Definitions of success and failure, including whether a valid call, successful tool execution, task completion, or recovery after an error counted as success.

These distinctions matter because an agent can produce arguments that match a schema but choose the wrong tool; it can call the right tool with valid arguments and still fail during execution; or it can encounter a validation error and recover successfully. A single pass/fail score would obscure where the behavior differed.

How documented tool-schema behavior differs

The documentation below describes particular framework surfaces, not a universal rule for every version or implementation. The available sources do not establish that any one of these behaviors caused a failure in the claimed runs.

Framework and documented scope Schema or tool behavior What that means for an evaluation
Microsoft Agent Framework, Go documentation A Go function signature determines its function tool’s input schema. The documentation also demonstrates structured fields with descriptions and enum constraints. Microsoft Learn: Function tools Record the function signature and the schema it produces; test whether the model’s arguments satisfy the declared fields and constraints.
OpenAI Agents SDK, JavaScript guide Standard Schema parameters are converted to JSON Schema and validated locally. For non-strict tools defined with raw JSON Schema, the developer is responsible for validating input. OpenAI Agents SDK: Tools Record which parameter-definition path is used and where validation occurs. A framework-generated validation error and an application-side validation error are different events.
AG2 tool documentation Only one tool per name is exposed in a turn; a later tool with the same name replaces an earlier one. This is a documented AG2 behavior. The cited page does not state a comparable argument-validation rule. AG2: Tools Check tool names and registration order as well as argument schemas. Do not assume this name-collision behavior applies to other frameworks.

A schema comparison should keep these layers separate: how a schema is declared, how it is presented to the model, whether arguments are validated, and what happens after validation. Documentation for one SDK or language does not by itself prove identical behavior elsewhere.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What a schema change can break

Without the experiment’s four actual variants, no specific change can be attributed to its runs. In general, these are distinct contract changes worth testing because each can expose a different mismatch between the model-facing description, validator, and tool implementation:

  • Adding a required field: calls that omit it may become invalid, even if earlier calls were accepted.
  • Changing a field’s type or name: generated arguments may no longer match what the validator or function expects.
  • Narrowing an allowed-value set: an argument that was previously permitted may become invalid.
  • Changing optionality or removing a field: callers and implementations may disagree about whether the field should be supplied or consumed.

These are examples of test dimensions, not a claim that they were the four changes in the headline. A useful test records the exact before-and-after schema and checks both input validity and downstream behavior.

Classify the failure before blaming the framework

Microsoft Research’s AgentRx taxonomy explicitly separates invalid tool calls from other agent failures, defining the category as “Invalid Invocation | Tool call malformed / missing args / schema-invalid.” Its report also distinguishes planning, intent, tool-output interpretation, guardrail, and system failures. Microsoft Research, March 12, 2026

That distinction is practical: a bad argument points toward a different debugging path than a correct tool call that returns an error or an agent that fails to use a useful result. AgentRx reports 115 manually annotated failed trajectories and an absolute improvement of 23.6% in failure-localization accuracy and 22.9% in root-cause attribution. Those figures belong to AgentRx’s own report; they are not findings from the claimed 36-run comparison.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For a schema-change evaluation, log each stage separately:

  1. Selection: Did the agent choose and invoke the intended tool?
  2. Argument conformity: Did the submitted input satisfy the declared schema?
  3. Execution: Did the function run, and did it return an error?
  4. Task outcome: Did the agent complete the requested task using the tool result appropriately?
  5. Recovery: If validation or execution failed, did the agent correct its call and continue?
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What related schema research can—and cannot—tell you

A separate May 4, 2026 arXiv preprint by Furkan Sakizli, titled TSCG, reports approximately 19,000 calls across 12 models and five scenarios and examines schema representations. TSCG preprint on arXiv Its scale and subject make it relevant background, but its benchmark is not evidence for the title’s 36 runs. The work is a preprint, and its results should be attributed to that study rather than treated as a replication of an undocumented experiment.

How to test tool-schema changes in production

Handle a tool schema as an API contract: make changes explicit, validate inputs at a known boundary, and keep regression results tied to the exact schema and framework configuration tested. A small repeated test can reveal a regression, but it should not be presented as a broad framework ranking unless its design and uncertainty support that conclusion.

  1. Save the contract. Keep the tool definition and its schema under version control. Record framework, SDK, model, and tool implementation versions alongside each test result.
  2. Validate at the boundary. Ensure every tool invocation is checked against the intended schema before the function runs. With the OpenAI Agents SDK JavaScript guide’s non-strict raw JSON Schema path, the documentation assigns input validation to the developer; other paths may document different validation behavior.
  3. Test old and new contracts. Include representative valid calls and invalid cases for each changed constraint. For a breaking change, test that rejected calls fail clearly rather than reaching the function in an unintended form.
  4. Repeat comparable trials. Keep the task, prompt, model settings, tool implementation, and run conditions constant when comparing schema versions. Report the run allocation and outcomes for each version, not just a total count.
  5. Keep outcome measures separate. Report argument validity, intended tool selection, execution success, task completion, and recovery independently. Include failure examples and explain how each category was counted.
  6. Check registration too. In frameworks where tool names and registration order affect exposure, test those separately from argument validation; AG2 documents one such same-name replacement behavior.

For a claim that one framework is more robust, the comparison should identify all configurations and publish enough detail to reproduce the task and classification. A count of runs without the allocation, outcomes, and uncertainty does not establish how generally the result applies.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.