iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
MCP resources, tools, and prompts are three different interfaces a server can expose to an AI application. Each one answers a different question: what information is available, what actions can be taken, and which reusable interaction pattern should start a task. None of them reduces tokens automatically. The token count an agent sends to a model depends on what the client loads, when it loads it, and what the model receives in each request.
The headline in this article describes a drop from 114K to 27K tokens in one author’s agent. That figure is a single reported result. Its model, counting method, and before-and-after configuration are not published, so it cannot be verified or generalized. The sections below explain the three primitives, what drives token use, and how to measure a change for yourself.
The three primitives at a glance
The Model Context Protocol (MCP) defines three server-side primitives. The table below summarizes how each one is typically used. The metaphors in the last column are explanatory shorthand, not formal protocol terms.
| Primitive | Practical role | Typical interaction | Who decides when it is used |
|---|---|---|---|
| Resources | Contextual data such as file contents, database records, or API responses | The client discovers and reads resources; the application decides how to use the data | The application or user |
| Tools | Actions or retrieval operations such as querying a database, calling an API, or computing a value | The model can see tool metadata and request a call; the client and server handle execution | The model requests, subject to client controls |
| Prompts | Reusable templates or instructions, optionally with arguments and examples | A user or application selects a named prompt and receives messages | The user or application |
The MCP architecture documentation illustrates the three together with a database server: a query tool, a schema resource, and a prompt containing few-shot examples. The architecture page opens by describing the protocol as having two layers, a data layer and a transport layer. It is the 2026-07-28 revision of that page. The resource, prompt, and tool specification pages are versioned 2025-06-18, so check which protocol revision your implementation uses before copying normative details.
#1 Best Overall
Resources: the context the application can read
A resource is data the server makes available for reading. It might be a file, a table schema, a record, or the output of an API endpoint. The important point for token use is that a resource does not enter the model’s context simply because it exists. The application chooses which resources to read and how much of each one to include.
That gives you control, but it also means the application carries the cost of that choice. Pulling in a full document when a short summary would answer the question is the most common way a resource inflates a request.
Tools: the functions the model can ask to run
The MCP tools specification says that servers can expose tools that language models invoke. Each tool has a name, a description, and an input schema. The model sees this metadata and can request a call with arguments. Tool results can include text, structured content, and resource links.
Do these 3 things before closing this tab:
1Scan for outdated or missing drivers - takes under a minute2Repair Windows errors before they cause bigger problems3Fix the driver behind crashes, sound loss and screen glitchesRank #2
Tools are where most of the per-request overhead tends to sit. Every tool definition the client sends becomes part of the input. A large set of tools, each with long descriptions and detailed schemas, adds tokens to every model call whether or not the model uses them. Tools also carry side effects, which is covered in the security section below.
Prompts: packaged starting patterns
A prompt is a named, reusable template. It can take arguments and return messages, such as a system-style instruction with few-shot examples. Prompts are chosen by a user or the application rather than by the model. Their token cost is incurred when they are selected, so a prompt is most useful when it replaces many ad hoc instructions, not when it is attached to every request by default.
Why a primitive is not a token-saving layer
Choosing resources over tools, or prompts over inline instructions, does not by itself lower the number of tokens sent to a model. Three factors determine the count:
- What enters the request. Tool definitions, schemas, resource contents, messages, and prior conversation turns all count.
- When it enters. Loading every tool at the start of a session differs from loading a relevant subset at the point of need.
- What the model receives. The same text can produce different token counts depending on the model, its encoding, and the language.
The official MCP documents describe what the primitives are for. They do not promise a reduction in tokens, and none of the reviewed sources report a fixed percentage saving from adopting one primitive over another.
Free tools Windows power users keep installed
One-click scans. No signup required.
What the 114K-to-27K figure does and does not show
The reported drop is 87,000 tokens. If both counts measure the same thing under the same conditions, that is about 76% of the original 114,000. That arithmetic is only meaningful if the two numbers are comparable, and the reported result does not say whether they are.
Several things would need to be known before the figure could be used as evidence:
Rank #4
- The model, its exact version, and the tokenizer or API endpoint used to count tokens.
- Whether the numbers are full input tokens, tool-definition tokens only, or cumulative conversation tokens.
- Whether messages, schemas, resource contents, output, cached input, and reasoning tokens are included.
- Whether the same user task and model settings were used before and after the change.
- Which of the three changes caused the difference. Isolating each one is the only way to attribute the saving to a specific layer.
- Whether the agent still completes the same tasks at comparable quality and latency.
Without those answers, the result describes one setup. It is not a protocol guarantee, and it should not be attributed to MCP, to any SDK, or to the model provider.
Reducing tokens without breaking tool use
The most concrete source of savings in the reviewed material is tool discovery. AWS Prescriptive Guidance on agent tooling explains that context use grows when every discovered tool is registered with the model. It gives an illustrative estimate of 250 to 500 tokens per typical tool definition, which would make 20 definitions cost roughly 5,000 to 10,000 tokens. These are AWS’s example figures, not measurements across MCP clients, and the actual size depends on schemas and serialization. The guidance recommends filtering or semantic search so the model sees a smaller, relevant subset.
Recommended Free Tools
The following practices follow from that reasoning. Treat each one as something to test in your own setup, not as a guaranteed outcome.
- Shorten tool descriptions carefully. Keep enough information for the model to select the right tool and build valid arguments. An aggressively shortened description can save tokens and still cause wrong calls.
- Expose a relevant subset of tools. Use filtering or runtime search where your client and server architecture supports it, so that a task about invoices does not carry the entire tool catalog.
- Keep metadata stable. A deterministic order for tool lists may help clients cache them. Caching is an implementation behavior, so do not assume it reduces billed input in every request.
- Read resources narrowly. If a URI, a summary, or a targeted query answers the task, avoid loading the full resource body. Confirm the saving by checking the actual request payload.
- Measure success alongside tokens. Track task completion, tool selection accuracy, and latency together with input usage. A token saving that lowers success is a regression.
How to measure input tokens accurately
Accurate measurement starts with the metric, not the tool. OpenAI’s Help Center article “Understanding and counting tokens” (updated 2026) explains that a plain-text count may leave out request structure, tools, schemas, images, and files. A useful comparison therefore counts the complete request that the model receives.
- Fix the model and settings. Use the same model, version, and generation settings for both runs.
- Fix the task. Run the same user request, ideally several times, so that variation in agent behavior is visible.
- Record the API usage fields. Use the input token figures the endpoint reports, which vary by provider and endpoint. Do not estimate from character counts.
- Separate the components. Report full request input, tool-definition tokens, and cumulative conversation tokens as distinct numbers.
- Note how cached and reasoning tokens are handled. State whether cached input and reasoning tokens are included in each figure.
- Change one layer at a time. Apply the resource, tool, and prompt changes separately, then together, so that each effect is visible.
Cost against quality: what the tool-description study found
A 2026 preprint, Model Context Protocol (MCP) Tool Descriptions Are Smelly!, analyzed 856 tools across 103 MCP servers. It reported that 97.1% of the descriptions had at least one identified quality issue, and that 56% did not state the tool’s purpose clearly. These figures describe that paper’s collected sample and scoring method.
The same study tested augmenting descriptions and compact variants. Full description augmentation produced a median task-success improvement of 5.85 percentage points and a 15.12% improvement in partial-goal completion. The same augmentation increased execution steps by 67.46% and caused regressions in 16.67% of cases. Compact variants reduced token overhead. The findings are specific to that study’s agents and tasks, so they show a trade-off to measure, not a rule for every agent.
Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minutePC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Security and consent when tools act
Token savings should never override control over what a tool can do. The MCP tools specification calls for a human to be able to deny tool invocations and for clients to give users clear signals before operations are carried out. A tool that writes data, sends messages, or calls a paid API needs confirmation logic that does not depend on how short its description is. Reducing token use by trimming tool definitions is only an improvement if these controls remain in place.
Quick Recap
*** ENDS ***
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

