Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Repair Windows errors before they cause bigger problems3Scan for outdated or missing drivers - takes under a minuteiTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
If your Claude Code usage is climbing faster than expected, check retained conversation context, extended thinking, MCP tools and results, and subagents or agent teams. Anthropic’s guidance identifies these as ways token use can accumulate; it does not establish that they affect every user or how much each contributes in a particular session. Start with /usage and /context, then change the workflow that matches what you find.
Find what is using tokens
In Claude Code, run /usage to inspect usage for the current session and /context to see what is occupying the context window. These commands answer different questions: usage is a starting point for token consumption, while context helps identify material being carried into requests.
Anthropic’s cost guide says that eligible Pro, Max, Team, and Enterprise plans also have a recent-usage breakdown that can attribute usage to skills, subagents, plugins, and individual MCP servers, along with behavior indicators such as long context and cache misses. Availability and attribution details can vary by Claude Code version, so check the version installed before relying on a particular display. The figures are approximate and calculated from local session history on that machine; they do not include activity on other devices or Claude.ai. See Anthropic’s Claude Code cost guide.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
For organization-wide monitoring, Anthropic documents the claude_code.token.usage and claude_code.cost.usage metrics, which can be grouped by dimensions such as token type, user, team, model, skill, plugin, or agent. Anthropic cautions that “Cost metrics are approximations.” Use the API provider’s billing records—such as Claude Console or the applicable cloud provider—as the source of truth for billed usage. See Anthropic’s monitoring documentation.
#1 Best Overall
Compare like with like when reading dashboards: Claude Code’s exported input count excludes cache reads and writes unless the cache token fields are added. Monitoring reports cache-read and cache-creation fields separately, so a total that omits them may not match a provider-side figure.
Four places tokens can accumulate
1. Old conversation context
Claude Code can carry earlier conversation context into later requests. Anthropic explains that “Token costs scale with context size: the more context Claude processes, the more tokens you use.” Unrelated follow-up work in a long session can therefore include context that no longer helps with the task.
Rank #2
When switching to unrelated work, use /clear to start with a clean conversation. If continuity matters, /compact can summarize the conversation instead. Compaction is not free: Claude must read and summarize the prior conversation. Anthropic recommends giving custom compaction instructions so the summary preserves only information needed for the next task. The cost and continuity trade-off is straightforward: clearing discards the session context, while compaction spends tokens to retain a smaller, selected version.
2. Extended thinking
Thinking tokens are billed as output tokens. Anthropic says the default thinking budget can reach tens of thousands of tokens per request, depending on the model. That can be disproportionate for a routine task that does not benefit from extensive reasoning.
Rank #3
Where the current model and task allow it, choose a lower effort level or disable thinking. Do not assume one setting applies to every model: Anthropic’s documentation distinguishes model behavior and releases, and some models always use extended thinking. Check the current model-specific controls in the cost guide before changing settings.
3. MCP servers and tool results
Configured MCP servers can add tool-definition context, and the tools’ returned content can add more. Anthropic says MCP tool definitions are deferred by default, so the overhead is not identical for every setup; large or verbose tool results can still consume context when used.
Rank #4
Use /context to identify context consumers and /mcp to review configured servers. Disable servers you are not actively using, and consider a CLI tool when it can accomplish the same task with less context. Keep results focused where possible: a tool that returns a large body of irrelevant output can cost more than its definition alone suggests.
Recommended Free Tools
4. Subagents and agent teams
Delegation adds requests rather than making their usage disappear. Subagents issue their own requests; they can keep verbose work out of the main conversation, but still contribute usage. Agent teams run separate instances, with each teammate using its own context window. Anthropic says team token use scales with the number of active teammates and how long they run.
Best Value
Anthropic gives an approximate comparison of about 7× the tokens of standard sessions when agent-team teammates run in plan mode. That figure applies to the described agent-team plan-mode case, not to every team or subagent workflow. Keep teams small, scope prompts tightly, use a lower-cost model for simple work where appropriate, and stop teammates once their tasks are finished. The trade-off is parallelism and time saved versus the extra requests and contexts it creates.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Choose the smallest useful change
| What you find | Try | Trade-off |
|---|---|---|
| Unrelated work in a long session | Use /clear; use /compact only when selected context is worth retaining. |
Clearing loses continuity; compaction uses tokens to summarize. |
| Thinking on a routine task | Lower effort or disable thinking if the model supports it. | Less reasoning may be unsuitable for complex work; model controls differ. |
| Unneeded MCP configuration or oversized results | Review with /mcp and /context; disable unused servers or narrow outputs. |
Removing integrations may make a workflow less convenient. |
| More agents than the task needs | Reduce the number of teammates, narrow assignments, and stop them when done. | Less parallel work may take longer, but creates fewer additional requests. |
When local usage does not match the bill
Local session estimates and exported metrics have scope and accounting limitations, so validate them against the billing records for the provider handling the requests. Routing also matters: a gateway can attribute and bill requests to the owner of the gateway credential on a per-token basis, and such requests may not count toward subscription usage limits. Anthropic says it does not endorse, maintain, or audit third-party gateways. If a gateway is involved, verify which credential is making the requests and where those requests appear in billing. See Anthropic’s gateway documentation.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

