tuesday, october 6, 2026 · the day's ai, attributed published by trilot llc · wyoming
guide · running the business

Routing tasks between AI models

Decide which jobs go to a cheap model and which stay on a frontier one, and move work between them without paying for the switch twice.

Published 2026-09-05 · Updated 2026-09-05 · Read 9 min · Reviewed by Rami Steitieh

Verified 2026-09-05 · Rami
on this page · 0 / 0 checked

You already route tasks between AI models. You just do it by which tab is open. The assistant you signed up for first became the default, and now everything goes there: the contract you need read carefully, and also the list of 40 addresses you want turned into a table. One of those jobs needs the model you are paying for. The other would have come back faster, and for a tenth of the money, from a model you already have access to.

This guide is the manual version of a problem large engineering teams solve with a trained classifier. If you run millions of requests a day, you will eventually build or buy one. If you run a business of one to 50 people, your routing lives in three places: a model picker in a chat window, a dropdown on an automation step, and a habit. It is not a guide to building routing infrastructure, and it will not tell you which model is best, because that answer expires. Every price below was read from the vendor’s own page on 5 September 2026, and vendors move these without much notice.

The spread between the cheapest and the dearest model is the entire reason to bother

On Anthropic’s price list, Haiku 4.5 is $1 per million input tokens and $5 per million output, Sonnet 5 is $2 and $10, Opus 5 is $5 and $25, and Fable 5.1 is $10 and $50 [1]. That is a factor of 10 across those four, on both halves of the bill. OpenAI’s current line is wider still: GPT-5.6 Luna at $0.20 input and $1.20 output against GPT-5.6 Sol at $5 and $30 is a factor of 25 either way [4]. Google prices Gemini 3.8 Flash at $0.75 input and $3.75 output on the paid tier through 31 December 2026, and free of charge on the free tier [6].

A 10x or 25x spread is a large number attached to a decision that costs you one dropdown. That is why routing gets attention out of proportion to how hard it is: the choice is cheap to make and the difference on the metered half of your bill is large.

There is a catch worth knowing before you run that arithmetic on your own spending. If you work in a chat seat rather than through an API, the spread is not money. A seat is a flat monthly fee, so routing a task down a tier buys you speed and spends your included usage more slowly. Those are real goods, particularly at three in the afternoon when you are near a limit, but they do not appear on an invoice. Check which currency you are actually saving before you spend an afternoon on this.

Route by how a mistake behaves, not by how important the job feels

Left to instinct, defaults drift upward, because the model you reach for is the one that matches how much the work matters to you. Importance and difficulty are different variables. A client’s invoice matters enormously and is trivial to extract fields from. A throwaway internal question can be genuinely hard.

Anthropic’s own guidance is a reasonable starting shape: choose Haiku for simple tasks, Sonnet for most production workloads, and Opus for the most complex reasoning [1]. That tells you the ladder exists. It does not tell you which rung your job sits on, and the vendor cannot, because it depends on what happens after the answer arrives.

The test that holds up is about the failure, not the task. Work belongs on a cheap model when the input is concrete and complete, so the model is not being asked to supply context it never got; when the set of acceptable outputs is small enough that you can tell a good one from a bad one at a glance; and when a wrong answer gets caught before it reaches anybody and costs you one retry. Reformatting a list into a table, pulling fields out of a receipt, tagging inbound email by topic, turning a transcript into notes, drafting a product description you were going to rewrite anyway, checking whether a document mentions a term: all of these pass, and all of them are the bulk of what most people actually run through an assistant.

Work stays on the expensive model when you cannot check the answer yourself, when the output goes to someone else without you reading it first, when the input is long and messy enough that the model has to hold a lot at once, and when a wrong answer would be plausible rather than obvious. That last one is the real criterion, and it is the one people skip. A cheap model rarely fails by producing nonsense you would spot. It fails by producing something reasonable and slightly wrong.

Turn the effort dial before you change the model at all

The step most people miss sits between the two models rather than at either end. Anthropic’s model-selection documentation notes that several Claude models support an effort parameter that trades intelligence for latency and cost within a single model, and states plainly that tuning effort is often a better lever than switching models [2]. Changing effort keeps everything else identical: same context handling, same behaviour, same prompt. Changing models changes all of it at once.

The same page sets out two documented starting points, and they suit different people. The efficiency-first path is to begin implementation with Haiku 4.5, test your use case thoroughly, evaluate whether performance meets your requirements, and upgrade only if necessary for specific capability gaps; the documentation recommends this for initial prototyping, applications with tight latency requirements, cost-sensitive implementations and high-volume straightforward tasks [2]. The capability-first path is to implement with Opus 5, optimise your prompts for that model, evaluate, then consider increasing efficiency by lowering effort or downgrading models over time, moving to Fable 5.1 only if evals at the higher effort settings still fall short on demanding reasoning or long-horizon agentic work [2].

Both paths hinge on the same thing, which the documentation is blunt about: create benchmark tests specific to your use case, because having a good evaluation set is the most important step in the process [2]. For a solo operator that sounds heavier than it is. An evaluation set is ten real jobs you have already done, saved alongside the answers you actually shipped. Run the cheaper model on those ten and read the outputs next to what you sent. Twenty minutes, once, per job type. Anything less rigorous than that is you deciding which model feels smarter, which is exactly the instinct that put everything on the expensive one.

The handoff is where the saving leaks back out

Moving a task between models mid-workflow costs something, and four things in particular do not travel.

Context is the first. A conversation’s history lives inside the tool that hosted it. Hand a half-finished task to a second model and you are re-pasting by hand, which is where people quietly give the new model less than the old one had, then conclude the cheaper model is worse.

Size is the second. Fable 5.1, Opus 5 and Sonnet 5 all list a 1M token context window with 128K maximum output; Haiku 4.5 lists 200K context and 64K output [3]. A job that fits comfortably on the frontier model may not fit on the cheap one, and the failure arrives at the worst moment, which is when the input is unusually large.

Recency is the third, and it surprises people. Published knowledge cutoffs differ by more than a year across a single vendor’s current line: Fable 5.1 at June 2026, Opus 5 at May 2026, Sonnet 5 at January 2026 and Haiku 4.5 at February 2025 [3]. Route a question about anything that changed recently to the model at the bottom of that list and you get a fluent answer from an older world.

Caching is the fourth. Anthropic prices a cache hit at 10% of the standard input price and states that caching pays off after one cache read at the 5-minute duration or two reads at the hour [1]. A workflow that alternates between models on the same long prompt writes a cache repeatedly and rarely survives long enough to read it back. If you are routing at volume, group the cheap work together and run it as a block rather than interleaving it.

There is a fifth cost that is not about performance. Sending the same material to a second vendor creates a second copy under a second retention policy. OpenAI states that by default, data from its business products and the API Platform after 1 March 2023 is not used to train its models unless you have explicitly opted in to share it, and that OpenAI may securely retain API inputs and outputs for up to 30 days to provide the services and to identify abuse, after which they are removed unless it is legally required to retain them [5]. That is not an argument against routing. It is an argument for knowing, before you paste, which vendors are holding your client’s document.

calculator
Monthly saving from routing
— $ / month

Defaults compare Claude Opus 5 output at $25 against Haiku 4.5 at $5. Output tokens only, because output is the dearer half on every list here. If you are on a flat seat rather than an API key, the answer is zero and the saving is time instead. Computed in the page; nothing is sent anywhere.

Write the mapping as tiers, and keep the model a setting you can change

Whatever you decide, record it as a list of jobs mapped to price tiers rather than to model names. Prices move in both directions rather than only downward, and the vendors publish the increases in advance: Google’s own pricing page schedules Gemini 3.8 Flash to go from $0.75 input and $3.75 output to $1.50 and $7.50 on 1 January 2027 [6].

Names churn as well, and old names keep their old prices. On that same Google page, Gemini 3.8 Flash sits at $0.75 and $3.75 while Gemini 3.5 Flash, described there as the earlier Flash model, is still listed at $1.50 and $9.00 [6]. An automation pinned to the older identifier is paying twice the input price and more than twice the output price for the same class of work. Pinning a version is the right call for reproducibility. Forgetting about it for a year is not.

That also means the model should be a field rather than a rebuild. Zapier’s own pitch for this is to use the models that work best for you and switch as things evolve, without rebuilding your workflows or changing how everything runs [8]. Whether that holds for your particular setup is worth testing on one low-stakes workflow before you rely on it, because a model swap can quietly change output formatting in ways that break the step downstream. Change the model on one automation, run it on live data for a week, and look at what the next step received.

A shipped router will take the routine volume, if you keep the override

Routing is increasingly something the tool does rather than something you do. Cursor Router analyses each request on query, context, task complexity and domain, combined with what Cursor knows about each model’s behaviour [7]. Cursor says it trained the router on more than 600,000 live requests and evaluated it in an online A/B test across millions of live requests, optimising for user satisfaction as the reward [7]. It ships three modes, Intelligence, Balance and Cost, described respectively as frontier quality, strong quality matching the models most people daily drive, and good quality while optimising token spend [7].

The published figures are specific enough to be useful. Cursor reports a cost per commit of $4.63 in Balance mode and $6.76 in Intelligence mode, against $7.34 for Opus 4.8 and $12.69 for Fable 5 [7]. It also reports that three high-volume accounts with thousands of users saved 30% to 50% on auto-routed requests versus routing everything to Opus 4.8, with no decrease in quality [7]. Read that second figure for what it is: three accounts, reported by the vendor whose router produced the saving. Your task mix is not their task mix, and the number you should care about is what happens on your work, over a week, next to whatever you were doing before.

The other cost of a shipped router is legibility. When you pick the model, a bad output has one obvious variable to check. When a classifier picks it, you have two, and you usually cannot see which model answered. That is a fine trade for the routine bulk of requests and a poor one for the small set of work that is expensive to unwind. Run the router on the first category and keep manual selection on the second.

checklist
Before you route a job to a cheaper model
0 of 8 · saved in this browser only

What still goes wrong

Every figure here came off a vendor page on 5 September 2026, and model names, prices and context windows all move. The structure outlasts the numbers. Check the failure rather than the importance, turn the effort dial before you change models, evaluate on ten real jobs, and write the mapping as tiers. Re-read the prices when you re-read the mapping.

The failure mode of routing is quieter than the failure mode of overspending, which is why it deserves more suspicion than it usually gets. A frontier model that struggles tends to say so or produce something visibly off. A cheaper model handed work you misclassified produces something fluent, well formatted and subtly wrong, and it does that reliably enough that a fortnight can pass before you notice the summaries have been drifting. If you route work down a tier and simultaneously stop reading the output, you have not saved money. You have moved the cost somewhere it will surface later, attached to a client rather than an invoice.

The last limit is arithmetic. Routing pays in proportion to volume, and most solo operators do not have volume. If you make 20 model calls a day inside a flat monthly seat, the saving available is zero dollars and some seconds, and the hour you spend building an evaluation set is an hour. The work is worth doing when a specific repetitive job runs hundreds of times a month through an API key, or when a metered tool has started producing a bill you cannot explain. For everything else, the honest answer is to keep using the good model and spend the attention on something that pays better.

sources
  1. 01Claude Docs — Pricingplatform.claude.com
  2. 02Claude Docs — Choosing a modelplatform.claude.com
  3. 03Claude Docs — Models overviewplatform.claude.com
  4. 04OpenAI — API pricingopenai.com
  5. 05OpenAI — Enterprise privacyopenai.com
  6. 06Google — Gemini API pricingai.google.dev
  7. 07Cursor — Cursor Routercursor.com
  8. 08Zapier — AIzapier.com
next guide
The org chart is not the signal
9 min · verified 2026-09-05
related guides