Free tools Windows power users keep installed
One-click scans. No signup required.
iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
In a small experiment described by the developer known as matsumotory, three AI coding agents called a custom MCP tool more often when its description included one sentence saying when to use it. Across 24 runs, the tool was called in 7 of the 12 runs whose description carried that sentence and in none of the 12 runs without it. That is a useful signal for anyone writing MCP tools, but the test was small and tightly scoped, and the author says it does not establish a general rule about how agents choose tools. The original write-up, published on DEV Community on September 29, 2026, is here.
What the experiment tested
The tool under test was a custom MCP server tool that hands several independent tasks to subagents so they run in parallel. The prompt given to the agents did not name the tool, so any call had to come from the agent noticing the tool in its available list and deciding to use it.
The author compared three agents, OpenHands, OpenCode, and Qwen Code, and varied two things:
The Tool Desk
Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →- Description guidance: the tool description either contained a sentence stating when to use the tool, or it did not.
- Rule file: each agent either received a startup rule file or started without one.
Those choices produce six agent-and-rule-file combinations. Each combination was run twice with the when-to-use sentence and twice without it, which gives the 24 runs.
#1 Best Overall
The results
| Description condition | Rule file | Runs | Runs with at least one tool call |
|---|---|---|---|
| Includes when-to-use sentence | Present | 6 | 5 |
| Includes when-to-use sentence | Absent | 6 | 2 |
| No when-to-use sentence | Present | 6 | 0 |
| No when-to-use sentence | Absent | 6 | 0 |
Taken together, the sentence condition accounts for all seven tool calls. The sentence did not need a rule file to produce calls, although it produced more of them with one. Without the sentence, the rule file produced no calls in this set of runs.
What the result does not establish
The author is clear that this is not a clean causal finding, and several limits follow from the design.
Rank #2
- Sample size. Each cell contains two runs per agent, so a single run changing outcome would shift the picture noticeably.
- Confounded comparison. The author notes that the main experiment and a separate rule-file test differ in the tool, the quality of the description, the task, and how specific the rule-file instructions were. The two cannot be merged into one comparison.
- A withdrawn explanation. The author first proposed that the weight of the agent’s processing decided whether descriptions or rule files mattered most. That conclusion was challenged during review and withdrawn. The idea remains a hypothesis: the evidence is consistent with it, but not enough to rule out effects from the rule file or to separate the confounding factors.
- A measurement defect. One agent’s subagent results were affected by a fault in the measurement program, as described below.
A stronger follow-up would vary description guidance and rule-file instructions in a full factorial design, while keeping the tool, the task, the agent versions, and the outcome measure fixed.
Recommended Free Tools
A tool call is not a completed task
The author separated two outcomes: whether the agent called the tool, and whether the subagents it launched actually finished their work. Four of the seven calls led to subagents completing their tasks. The results varied by agent:
- Qwen Code completed subagent tasks in two of its three calls. In a further call, seven of eight subagent tasks completed.
- OpenCode completed a run only after the author corrected an error in its launch script.
- OpenHands made calls, but none of the 32 subagents it launched completed. The author attributes this to a defect in the measurement program rather than to the agent’s behavior.
Across the runs, the author reports that all eight tasks passed their tests and that the agents did not rewrite the tests to make them pass.
Rule files: a separate consultation-tool test
A second, distinct experiment used a consultation tool rather than the parallel-execution tool. When the rule file included a provision telling the agent to consult the tool, the agent called it in all six runs. Without that provision, there were no calls in six runs. Because the tool and task differ from the main test, these results cannot be compared directly with the 24-run comparison. They do suggest that explicit rule-file instructions can drive tool use in their own right, which is a separate question from whether a description sentence does.
Rank #4
What related studies do and do not show
Two 2026 papers on tool descriptions are often cited alongside this kind of experiment. The author says they measured task success, not whether an agent decided to call a tool, so they should not be read as direct evidence for the call-frequency result. The figures below are as the article reports them.
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Repair Windows errors before they cause bigger problems3Scan for outdated or missing drivers - takes under a minute- Hasan and colleagues, arXiv 2602.14878 (“MCP Tool Descriptions Are Smelly! Towards Improving AI Agent Efficiency with Augmented MCP Tool Descriptions”; submitted 16 February 2026, revised 31 May 2026). The sample covered 856 MCP tools from 103 MCP servers. The article reports that 97.1% of sampled descriptions had at least one defect and that 56% did not state their purpose explicitly. Augmenting descriptions produced a median 5.85 percentage-point increase in task success. Execution steps rose by 67.46%, and performance declined in 16.67% of cases.
- Guo and colleagues, arXiv 2602.20426 (“Learning to Rewrite Tool Descriptions for Reliable LLM-Agent Tool Use”; submitted 23 February 2026, revised 29 April 2026). In an experiment with at least 150 candidate tools, rewritten descriptions produced an average 60.89% improvement in per-query success compared with the original descriptions.
What the specification says
According to the author’s reading of the MCP 2026-07-28 specification, a tool’s description is defined as a human-readable account of what the tool does. The specification does not require that it also say when the tool should be used. A when-to-use sentence is therefore an addition you make to the description, not a field the protocol asks for.
Best Value
Applying the finding in your own setup
If you want to test this with your own tools, the author’s design points to a practical sequence:
- Add a single when-to-use sentence that names the situation, for example: “Use this tool when a task splits into independent subtasks that can run in parallel.”
- Keep the rule-file guidance separate, and record which of the two you changed in each run.
- Log the exact description each agent received, so you can confirm what the agent saw when it made its decision.
- Count tool calls and completed work as separate measures, and check each subagent or downstream result before counting it as a success.
- Change one variable at a time, or cross the description and rule-file variables in a 2×2 grid, and hold the agent version, task, and tool constant.
- Run more than two repetitions per cell before drawing a conclusion from any difference.
The author summarizes the change in approach this way: “I now write the results of the two measurements separately.” The measurements in question are tool calls and subagent completions, and the proposed explanation for the difference remains, in the author’s words, a hypothesis.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

