When to split a job across parallel agents
Decide whether a task is worth fanning out to several agents at once, what the fan-out costs, and which jobs get worse when you split them.
on this page · 0 / 0 checked
Every so often a lab publishes something hard and the write-up mentions, almost in passing, what the search cost to run. OpenAI’s post on ten results in mathematics and theoretical computer science credits them to “an internal version of Astra, our next major model”, says that afterward “the model formalized each argument in a Lean certificate”, and reports that the tokens needed to find the solutions “would cost roughly $2,000 at Sol API rates” [8]. You read that, then you look at the single agent grinding through your own work one step at a time, and a thought arrives: you should be running ten of these.
Sometimes you should. More often the second agent makes the job slower, more expensive and harder to check, and you find that out three weeks in, after the pattern has spread through your setup. There is a test that tells you which case you are in, and it takes about a minute to apply. This guide is for people who already run an agent that does real work, in Claude Code, the Agent SDK, the OpenAI Agents SDK, or a scripted loop of API calls, and are deciding whether to fan it out. If you are using a chat window and want better answers, none of this applies to you; better prompts will do more.
One agent with more tools is the right default
Start from the position both major vendors publish. OpenAI’s guide to building agents is blunt about the order of operations: the recommendation is to “maximize a single agent’s capabilities first”, because “A single agent can handle many tasks by incrementally adding tools, keeping complexity manageable” [5]. Multiple agents, it says, “can introduce additional complexity and overhead, so often a single agent with tools is sufficient” [5]. Anthropic’s version is the same instruction in different words: “Start with simple prompts, optimize them with comprehensive evaluation, and add multi-step agentic systems only when simpler solutions fall short”, and add complexity “only when it demonstrably improves outcomes” [2].
That is not modesty. It is that the failure you are trying to fix is usually not a shortage of agents. OpenAI names two things that genuinely justify splitting. One is logic that has grown too tangled to keep in a single prompt, when the instructions fill with conditionals that no longer scale. The other is tool overload, and the interesting part is that the problem is similarity rather than count: “Some implementations successfully manage more than 15 well-defined, distinct tools while others struggle with fewer than 10 overlapping tools” [5]. If your agent keeps picking the wrong one of three search tools that do nearly the same thing, deleting two of them fixes more than adding an orchestrator will.
There is also a category error worth catching early. If you already know the steps and their order, you do not want agents at all. Anthropic draws the line clearly: workflows “offer predictability and consistency for well-defined tasks”, while agents are “the better option when flexibility and model-driven decision-making are needed at scale” [2]. A fixed sequence of five steps that runs the same way every time is a workflow, and in a Zapier or n8n automation the order lives in the tool rather than in a model’s judgement.
The test is whether the pieces need to talk to each other
Here is the minute-long test. Write down the subtasks. Then ask whether any of them needs to know what another one found in order to do its own work. If the answer is yes for more than one pair, stop.
Anthropic’s engineering post on its multi-agent research system says exactly where the pattern earns its keep: tasks “that involve heavy parallelization, information that exceeds single context windows, and interfacing with numerous complex tools”, and above all “breadth-first queries that involve pursuing multiple independent directions simultaneously” [1]. Fifteen suppliers to check against the same six criteria is that shape. Twenty-seven jurisdictions, each read the same way, is that shape. Each track can finish without ever hearing from the others.
The same post is equally clear about the other side, and it is the sentence most people skip: “Some domains that require all agents to share the same context or involve many dependencies between agents are not a good fit for multi-agent systems today. For instance, most coding tasks involve fewer truly parallelizable tasks than research” [1]. A vendor telling you its own architecture is wrong for the single most common agent use case is worth more than any benchmark on the page.
Two shapes qualify, and they are different jobs. Anthropic calls the first sectioning, “Breaking a task into independent subtasks run in parallel”, and the second voting, “Running the same task multiple times to get diverse outputs” [2]. Sectioning buys coverage. Voting buys confidence, and it is the one operators forget: running the same risky classification three times and comparing is a legitimate use of parallelism that requires no decomposition at all. Both are useful “when the divided subtasks can be parallelized for speed, or when multiple perspectives or attempts are needed for higher confidence results” [2]. When you cannot list the subtasks in advance, you are in orchestrator-workers territory, where “a central LLM dynamically breaks down tasks, delegates them to worker LLMs, and synthesizes their results” [2]. That flexibility is the whole point of the pattern and also the reason it is hard to budget.
The real gain is context, not speed
Most people reach for parallel agents to go faster. Speed is the smaller prize. The larger one is that each subagent gets a clean context window, and context is the scarce resource.
Anthropic’s context engineering post states the mechanism plainly: “As the number of tokens in the context window increases, the model’s ability to accurately recall information from that context decreases”, so context “must be treated as a finite resource with diminishing marginal returns”, drawn against what the post calls an “attention budget” [3]. An agent that has read forty files is worse at the forty-first than it was at the first. Splitting is how you stop paying that tax. In the sub-agent architecture the post describes, “Each subagent might explore extensively, using tens of thousands of tokens or more, but returns only a condensed, distilled summary of its work (often 1,000-2,000 tokens)” [3]. Thirty thousand tokens of searching become fifteen hundred tokens of finding, and only the finding lands in the main thread.
The Agent SDK documentation describes the same behaviour from the implementation side: “intermediate tool calls and results stay inside the subagent; only its final message returns to the parent”, so “A research-assistant subagent can explore dozens of files without any of that content accumulating in the main conversation” [4]. That is the benefit worth designing for. Faster is a bonus.
One mechanical detail causes most first attempts to fail, and it follows directly from the isolation. “The only content you pass from parent to subagent is the Agent tool’s prompt string, so include any file paths, error messages, or decisions the subagent needs directly in that prompt” [4]. A subagent does not inherit the conversation, the parent’s system prompt, or the tool results you were both looking at. If you brief it the way you would brief a colleague who was in the room, it will go and find the wrong thing very efficiently. The other benefit that comes free is restriction: a reviewer subagent given only read tools “can analyze but never accidentally modify” the files it reads [4], which is a cheaper safety mechanism than any instruction you could write.
Budget for roughly fifteen times the tokens
The cost is not a rounding error, and Anthropic publishes the multiplier. Agents “typically use about 4× more tokens than chat interactions”, and “multi-agent systems use about 15× more tokens than chats” [1]. Those figures come from Anthropic’s own research product rather than from your workload, but the order of magnitude is the thing to plan around: a fan-out is not 20% more expensive than a single agent, it is a few times more expensive.
You get something for it. In Anthropic’s internal research evaluation, “a multi-agent system with Claude Opus 4 as the lead agent and Claude Sonnet 4 subagents outperformed single-agent Claude Opus 4 by 90.2%”, and the analysis found that “token usage by itself explains 80% of the variance”, with the number of tool calls and the model choice as the other two factors [1]. Read that second clause carefully before you celebrate the first. If most of the gain is explained by spending more tokens, then a large part of what you bought with the swarm was the spending, and the architecture was mainly a way to spend it usefully. Anthropic reaches the honest conclusion itself: “For economic viability, multi-agent systems require tasks where the value of the task is high enough to pay for the increased performance” [1].
The lever that makes the arithmetic survivable is in the configuration above: a strong lead agent with cheaper workers. Current list prices make the gap concrete. Claude Opus 5 is $5 per million input tokens and $25 per million output; Claude Sonnet 5 is $2 and $10; Claude Haiku 4.5 is $1 and $5 [7]. Workers usually do the bounded, well-specified half of the job, which is the half that survives a cheaper model. Moving six subagents from Opus 5 to Sonnet 5 cuts their output cost by 60% before you change anything about the design [7]. The orchestrator is where the judgement lives, so that is where you spend.
tasks × subagents × output tokens × price. Default price is Claude Sonnet 5 output at $10 / MTok [7]. Input tokens and the orchestrator's own turns are excluded, so treat the result as a floor. Computed in the page; nothing is sent anywhere.
Put the plan in code when the steps are knowable
There are two ways to decide who does what, and the choice is a cost and predictability decision rather than a philosophical one. With model orchestration, “given an open-ended task, the LLM can autonomously plan how it will tackle the task, using tools to take actions and acquire data, and using handoffs to delegate tasks to sub-agents” [6]. With code orchestration, your own logic sets the sequence, which the OpenAI Agents SDK documentation says “makes tasks more deterministic and predictable, in terms of speed, cost and performance” [6]. If you can write your fan-out as a loop over fifteen suppliers, write it as a loop. Independent work runs concurrently through ordinary language primitives such as asyncio.gather [6], and a loop costs the same amount every Tuesday.
When the orchestration does belong to a model, pick the shape deliberately. Agents-as-tools, where a manager calls specialists and keeps the answer, suits the case where you want “one agent to own the final answer, combine outputs from multiple specialists”; handoffs, where a specialist takes over the conversation, suit the case where you want “the specialist to respond directly, keep prompts focused” [6]. OpenAI’s guide names these the manager pattern and the decentralized pattern, and puts the manager pattern as “ideal for workflows where you only want one agent to control workflow execution and have access to the user” [5]. For anything a client sees, that is usually the one you want, because a single agent owning the final output is also a single place to fix the tone.
Scale changes the answer again. Anthropic’s own documentation says subagents “work well for a few delegated tasks per turn”, and for runs coordinating dozens to hundreds of agents recommends moving “the orchestration into a script the runtime executes outside the conversation context” [4]. A conversation is a poor control loop for sixty parallel jobs, and the fix is not a better prompt.
Cap depth, concurrency and spend before the first run
The specific way this goes wrong is worth stating in the vendor’s own words: “Claude decides on its own when to spawn a subagent and how many to spawn. Each subagent makes its own API requests, which count toward the query’s total_cost_usd, and a subagent can spawn subagents of its own, so one prompt can grow into a tree of agents” [4]. Nothing about that is a malfunction. It is what delegation means when the delegate can also delegate.
Three limits bound it, and the defaults are worth knowing before you meet them. Nesting depth defaults to three layers of subagents below the main agent, and setting CLAUDE_CODE_MAX_SUBAGENT_SPAWN_DEPTH to 1 stops your subagents from spawning any of their own. Concurrency defaults to 20 subagents running at once through CLAUDE_CODE_MAX_CONCURRENT_SUBAGENTS, and beyond that a spawn attempt comes back with Concurrent subagent limit reached. Spend has no default limit at all; you set maxBudgetUsd in TypeScript or max_budget_usd in Python, and at the cap the run refuses to spawn more subagents, stops background subagents still running, and ends with the error_max_budget_usd result subtype [4]. Twenty concurrent agents three layers deep with no ceiling on spend is a plausible accident, not a worst case.
Model choice interacts with this. The documentation notes that “Claude Opus 5 delegates to subagents more readily than earlier models, so the depth, concurrency, and spend limits matter most on queries that run Opus 5” [4]. A prompt that behaved modestly on last year’s model can fan out considerably on this one, which is a good argument for setting the caps once, in your project configuration, rather than deciding each time.
What still goes wrong
Verification does not parallelise, and that is the constraint that eventually binds. The mathematics case is instructive precisely because it had an answer to this problem that you will not have: “the model formalized each argument in a Lean certificate” [8], which is a check software performs without anyone’s permission. Notice also that OpenAI records human help in the loop, saying “We helped prepare the manuscripts and formalize the proofs in Lean, and we take responsibility for their correctness” [8]. There is no Lean certificate for a supplier shortlist or a market scan. Six agents produce roughly six times the material for one human to read, and if you have not budgeted that reading time you have not made the job faster, you have moved the bottleneck onto yourself and hidden it.
The headline evidence for the pattern is narrower than it sounds. The 90.2% figure comes from Anthropic’s internal research eval, run on its own multi-agent system against its own single-agent baseline [1], and the same post says shared-context and dependency-heavy domains, including most coding, are a poor fit today [1]. It is a real result about a specific kind of work. It is not a general claim that more agents produce better outputs, and it does not transfer to your invoice reconciliation by analogy.
Debugging gets meaningfully worse, and this is the cost nobody prices. When one agent goes wrong you read one transcript. When a tree of agents goes wrong you read a tree of transcripts, most of them irrelevant, to find the subagent that was briefed badly three levels down. The isolation that saved your context window is the same isolation that hides the mistake. Keep the runs small enough that you would be willing to read all of them, and grow only when a real failure, on real work, tells you that one agent was not enough.
- 01Anthropic — How we built our multi-agent research systemanthropic.com
- 02Anthropic — Building effective agentsanthropic.com
- 03Anthropic — Effective context engineering for AI agentsanthropic.com
- 04Claude Agent SDK — Subagentscode.claude.com
- 05OpenAI — A practical guide to building agentscdn.openai.com
- 06OpenAI Agents SDK — Orchestrating multiple agentsopenai.github.io
- 07Anthropic — Model pricingplatform.claude.com
- 08OpenAI — Ten advances in mathematics and theoretical computer scienceopenai.com