Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

Sub-agents reduce engineering cost only under specific conditions. A coordinator that hands independent, tightly scoped pieces of work to worker agents can finish a large job sooner and keep each worker’s context small. The same design usually consumes more tokens in total, because the coordinator plans the split, each worker reads its own inputs, and someone has to merge and check what comes back. For a short task, a chain of steps where each one needs the previous result, or anything that already fits comfortably in one context, a single agent is the cheaper and simpler choice.

The useful question is therefore not whether sub-agents are good, but whether a particular workload splits cleanly enough for the added coordination to pay for itself. The sections below give a decision rule, the cost components to count, what the published vendor figures do and do not show, and a procedure for testing the trade-off on your own work.

When to delegate and when to keep one agent

Delegation earns its overhead when three conditions hold together: the pieces do not depend on each other’s results, each piece can be stated as one question or deliverable, and the combined material is large enough that splitting it either saves elapsed time or stops one agent from reading everything. OpenAI’s multi-agent guide draws the same line in two short rules. The first reads: “Use subagents for independent tasks, such as reviewing separate documents or investigating different causes of a failure.” The second reads: “Keep short tasks and dependent steps in the main agent.” (OpenAI, Agents API multi-agent guide)

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Anthropic’s cost guidance states the default more bluntly: “If the work is one chain, fits in one context without a long cost tail, or a single model at lower effort already meets your bar, don’t build an orchestrator.” (Anthropic, Claude platform cost-and-intelligence guidance)

Situation Recommended path Reason
Several independent files or documents to review, with separate outputs Delegate, one worker per package Packages can run in parallel, and each worker only reads its own slice
Several possible causes of one failure to investigate Delegate, one worker per hypothesis Hypotheses are independent; the coordinator compares the findings
Input larger than one practical context window Partition first, then delegate if the parts are independent Avoids one agent holding or re-reading material it cannot fit at once
A short task such as a single bug fix or a small refactor Single agent Planning and synthesis overhead can exceed the work itself
A dependent chain, such as design, then implementation, then testing Single agent, or serial hand-offs Parallel workers cannot shorten a chain where each step needs the previous output
Work that fits one context, where one model at lower effort meets the quality bar Single agent No additional coordination is needed to reach the same result
Routine work where most items are cheap but a few are costly Consider delegation, and measure before adopting it Vendor examples show savings in some configurations, but only as measured results on specific tasks

What counts as cost in a multi-agent run

A per-token price tells you what one model call costs. A multi-agent run has to be priced as a whole, so count every stage:

  • Coordinator planning. Tokens spent reading the task, deciding the split, and writing each worker’s instructions.
  • Worker context. Each worker begins with its own instructions, tool definitions and the material it must read. Shared background is paid again for every worker that receives it.
  • Tool calls. Searches, file reads and test runs made by each worker, including ones that turn out to be unnecessary.
  • Retries. Workers whose output is incomplete, off scope or wrong and must be rerun.
  • Synthesis. The coordinator reading every worker’s output, resolving disagreements and writing the final answer.
  • Human integration. Review and merge time after the run. It does not appear in token counts, but it often decides whether the approach saved anything.

Anthropic’s token multipliers

Anthropic reports from its own usage data that “In our data, agents typically use about 4× more tokens than chat interactions, and multi-agent systems use about 15× more tokens than chats.” The same post says the economics only work for tasks valuable enough to justify the performance gain. These figures come from Anthropic’s 2025 engineering post on its multi-agent research system; the page does not show an exact publication day. The multipliers describe chat-versus-agent usage in general, not coding tickets specifically. (Anthropic, multi-agent research system)

What the published figures show, and what they do not

Vendor results exist, but each one is tied to a specific benchmark, configuration and metric. The table lists them as the sources report them.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Source and date label Setup Reported result Limits stated by the source
Anthropic multi-agent research system post (2025; no exact day shown) Claude Opus 4 lead with Claude Sonnet 4 subagents, compared with single-agent Claude Opus 4 90.2% improvement on Anthropic’s internal research evaluation Internal evaluation; not a coding productivity guarantee
Anthropic cost guidance, corpus benchmark (2026 platform documentation; no exact date shown) One Claude Fable 5.1 lead and 25 Claude Sonnet 5 workers, on a 21.6-million-token corpus benchmark About 2.3 hours, against 15–20 hours for the solo configuration Vendor benchmark; the page describes a platform-reported limit, not ordinary engineering tickets
Same corpus benchmark, cost (same source) Same coordinator configuration, compared with the solo configuration 47%–55% lower cost; scores 10–12 points below the solo configuration The score gap is material; the saving is not a general one
DRACO test (same source) Same-model agents given time instructions and an elapsed-time clock 33% less elapsed time and 54% lower cost per task; 1.5-point lower score The clock was not measured with lower-cost workers, and coordinator-only clock visibility was not tested
BrowseComp slice (same source) Claude Fable 5 coordinator with one Claude Sonnet 5 worker, on a deliberately easy 10-problem slice About half the average cost, and roughly one-third the 90th-percentile cost ($12 against $33) Easy sample, not to be extended to harder traffic; the most expensive solo run cited cost $84 and gave a wrong answer

Three points follow from the table. First, the time gain and the cost gain came from different configurations, so one headline number does not describe both. Second, the cheaper configurations scored lower, and the sources do not show that this trade is acceptable for your work. Third, none of these figures measures everyday engineering tickets. No independent, cross-provider study of coding costs is available to cite here, so treat these numbers as vendor results for named benchmarks.

How to orchestrate sub-agents, step by step

Use this sequence for any workflow you intend to delegate. If a step shows the work is serial or fits one context, stop and run a single agent.

Step 1: Classify the task

Map the work before writing any prompt. List each work package, the files or inputs it touches, and what it depends on. Then answer three questions:

  • Can each package finish without another package’s output?
  • Do any two packages write to the same file or record?
  • Does the total input exceed what one agent can practically hold in one context, or does it only feel large?

If most packages fail the first question, keep the work serial in a single agent.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Step 2: Write a task contract for each worker

Each worker receives one question or deliverable, only the context and tools it needs, and a short description of the expected output. A usable contract for a review task looks like this:

  • Task: Review the input-validation logic in the files under src/auth/ for defects that could let untrusted input reach a database query.
  • Scope: Read-only. Do not edit files or review files outside the listed directory.
  • Tools: File read and text search only.
  • Output: A JSON array of findings, each with file, line, severity and one sentence of evidence. Return an empty array if nothing qualifies.

Avoid sending the same broad prompt to several workers unless diversity of approach is the goal. Identical prompts usually produce duplicated reading and duplicated spend.

Step 3: Set concurrency and stop conditions

Choose a ceiling on how many workers run at once. Give every worker a stop condition: a maximum number of tool calls, a time limit, or a rule to return partial results with a flag instead of continuing to search. Workers that touch shared files need a coordination rule. Either assign each file to one worker, or have workers return proposed changes and let the coordinator apply them in sequence.

Platform defaults for concurrency and related settings differ and change over time, including between stable and beta interfaces. Check the current settings in OpenAI’s Responses multi-agent documentation or Anthropic’s platform documentation before you hard-code a value.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Step 4: Synthesize and verify

Specialization narrows each worker’s prompt and tools, which is part of the cost argument. Parallel outputs are still not a finished answer. Anthropic’s Managed Agents documentation describes a coordinator and worker pattern in which each agent runs in an isolated context, and the coordinator reconciles the workers’ results into one answer (Anthropic Managed Agents, multi-agent orchestration). In practice the coordinator checks each finding against its evidence, resolves contradictions, confirms that the pieces fit together, and passes the result to tests or human review. Delegation does not remove those responsibilities; it adds more input to them.

Step 5: Measure the whole run against a single-agent baseline

Pick representative tasks from your own backlog, including some you expect to be poor fits for delegation. Run each task twice: once as a single agent and once orchestrated. Record the metrics below for both runs.

Metric What to record Where to get it
Total tokens Input and output tokens for every agent, coordinator included Usage reporting for each model call
Total cost Spend per task, including retries and synthesis Billing or usage logs
Elapsed time From task start to an accepted result, not to the first worker finishing Your own timestamps
Retries Number of reruns and the reason for each Run logs
Quality Defects found in review, test pass rate, reviewer acceptance Your test suite and review notes
Integration effort Minutes spent merging and checking worker output Your own time tracking

These are practical measurement choices drawn from how orchestration and cost work, not a published formula. The orchestrated run is worth keeping only if it beats the baseline on the metrics you care about, with quality not materially worse.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Implementation options

Two hosted platforms document this pattern directly. OpenAI’s Agents API multi-agent guide covers sub-agent delegation, and OpenAI’s Agents API overview describes managed sessions, orchestration, context compaction, recovery and delegation. Anthropic’s Managed Agents documentation describes a coordinator and worker pattern with isolated agent contexts. Both are documented, but model choices, pricing, feature names and beta status change, so confirm them on the current official pages before you budget a workflow.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

When costs rise instead of fall

If an orchestrated run costs more than the single-agent baseline without a clear time gain, the cause is usually one of the following.

Symptom Likely cause Fix
Token total above baseline, with no time saving Workers were dependent, waited on each other, or re-read the same material Re-check the dependency map from Step 1 and move the chain back to one agent
Workers return conflicting answers Overlapping scopes, or the same broad prompt sent to several workers Narrow each scope and assign disjoint files or sections
Coordinator context keeps growing Full worker transcripts are returned instead of summaries Require the fixed output format from Step 2 and pass only findings and evidence pointers
Frequent reruns Vague output expectations or no stop condition Tighten the task contract and cap retries as set in Step 3
Quality drops after moving workers to a cheaper model The subtask needed more reasoning than the cheaper model delivers Compare per-task quality and keep the stronger model for the hard subtasks
Edits collide in shared files Two workers changed the same file Apply the file-ownership rule from Step 3
Faster on paper, slower in practice Merge and review time was not counted Add integration effort to the Step 5 measurements

n

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.