How to route work to the cheapest model that still works
Sort your recurring AI tasks into cheap, mid and top lanes, so a free window closing or a price changing costs you an edit rather than a rebuild.
on this page · 0 / 0 checked
The offer shows up every few weeks in some form. A model is free on an aggregator for a fortnight. An open-weight release lands with a licence that lets you run it on your own machine. A free tier quietly gets a better default model. Each time, the same two mistakes are available. You can ignore all of it and keep running every task through the most expensive model you have access to, because that feels like the safe choice and the bill is only annoying rather than alarming. Or you can chase each announcement, rewire your work around whatever is free this month, and spend more hours migrating than you ever saved in fees.
There is a third option, and it is boring in the way that useful things usually are. You decide once which of your recurring tasks can run on a cheap model, which need a mid model, and which genuinely need the top one, and you set things up so moving a task between those lanes is an edit rather than a project. This guide is for a solo operator or small team with a handful of tasks they run weekly, paying for one or two AI products and wondering whether the money is going to the right place. It is not for anyone whose vendor is fixed by a procurement or compliance policy, and it is not worth the overhead if your entire AI usage is a few chat messages a day. Prices below are US list prices fetched on 4 September 2026.
”Free” describes four different arrangements
The word covers at least four things, and they fail in different ways.
The first is the free tier of a consumer chat product. Claude’s Free plan gives you Haiku and Sonnet on web, iOS, Android and desktop, and does not include Opus or Fable [1]. The cap is expressed as usage per session rather than as tokens: the same page says Pro gives you at least 5x more usage per 5-hour session than Free, without publishing a number for either [1]. That is a lane you can lean on and cannot plan against.
The second is a free tier on a developer API. Google’s Gemini API lists a long set of models as “Free of charge” on the free tier, including Gemini 3.8 Flash, Gemini 3.5 Flash-Lite and Gemini 2.5 Pro, while Gemini 3.1 Pro Preview is not available on the free tier at all [3]. This is the lane people mean when they say a frontier-class model is free to use programmatically.
The third is a free route on an aggregator, where a model appears with a :free suffix and no per-token charge. OpenRouter caps those routes at 20 requests per minute and 50 requests per day for accounts that have never purchased at least $10 in credits, rising to 1,000 requests per day for accounts that have [7]. It also notes that if your account has a negative credit balance you may see 402 Payment Required errors, including for free models [7].
The fourth is open weights. OpenAI released gpt-oss-120b and gpt-oss-20b on 5 August 2025 under the Apache 2.0 licence, with the 120b model designed to run on a single 80 GB GPU and the 20b model on edge devices with 16 GB of memory, both supporting context lengths up to 128k [8]. That is the only one of the four that nobody can reprice, because it is a file you already have.
The first three are pricing decisions that someone else revisits on their own schedule. Treating them as infrastructure is the error. Treating them as a discount you take while it lasts, on work you could move tomorrow, is the correct posture.
The bill for a free lane arrives as data rights and rate limits
The dollar price is not the whole price, and the difference is written down on pages most people never open.
Google is unusually direct about it. The Gemini API pricing page marks free tier usage with the note that content is used to improve Google products, and marks the paid tier as not used that way [3]. The terms go further. For the unpaid services, Google uses the content you submit and any generated responses to provide, improve, and develop Google products, human reviewers may read, annotate, and process your API input and output, and the terms instruct you plainly not to submit sensitive, confidential, or personal information to the unpaid services [4]. On the paid services, Google states that it does not use your prompts or responses to improve its products [4].
The consumer products split along a similar line, but not in the same direction at every vendor, which is why assuming is expensive. OpenAI says ChatGPT improves by further training on the conversations people have with it unless you opt out, and that you can opt out through the privacy portal by clicking “do not train on my content”; separately, it says that by default it does not train on any inputs or outputs from its products for business users, including the API [5]. Anthropic’s consumer privacy page describes the opposite default for chats, using them to improve Claude when you choose to allow it, and notes that Incognito chats are not used to improve Claude even when model improvement is enabled [6].
So the practical rule writes itself. Anything covered by a client NDA, anything containing someone else’s personal data, and anything you would not want a human reviewer reading goes down a paid lane or a local one. Everything else, which for most operators is the majority of the work, is fair game for the cheap lane. The second cost is throughput. A route capped at 50 requests a day [7] is fine for drafting and useless for a batch job, and finding that out at the point of use is how a Tuesday disappears.
Inside one vendor, the ladder is a clean multiple
Once you are paying, the choice gets easier to reason about, because the price ladders are tidy.
Anthropic’s API prices are Fable 5.1 at $10 per million input tokens and $50 per million output, Opus 5 at $5 and $25, Sonnet 5 at $2 and $10, and Haiku 4.5 at $1 and $5 [1]. Haiku is exactly a tenth of the frontier price on both sides, Sonnet exactly a fifth. OpenAI’s ladder is steeper. GPT-5.6 Sol costs $5 per million input tokens and $30 per million output, GPT-5.6 Terra costs $2 and $12, and GPT-5.6 Luna costs $0.20 and $1.20 [2]. Sol is exactly 25 times Luna on both sides. Google’s Gemini 3.8 Flash is $0.75 and $3.75 on the standard paid tier, a price the page marks as running through 31 December 2026 [3], which is a useful reminder that a published rate can carry an expiry date in the small print.
This is the shape that makes routing worth doing. You are not choosing between a good model and a bad one. You are choosing whether a specific piece of work is worth 25 times what it would cost one rung down, and for a surprising amount of routine output the honest answer is no.
Price one real workload before deciding anything. A single recurring task, run at a realistic volume, usually settles the argument faster than a comparison table.
Priced at GPT-5.6 Luna rates of $0.20 per million input tokens and $1.20 per million output tokens. GPT-5.6 Sol is exactly 25 times these figures on both sides, and GPT-5.6 Terra 10 times [2]. Computed in the page; nothing is sent anywhere.
Read the result against your own time, not against zero. If the cheap lane comes to $2 a month and the top lane to $50, the $48 is only worth spending where it buys an outcome you would otherwise have to fix by hand. If both numbers are under the price of a coffee, stop optimising and use the better model, because at that volume the routing logic costs more attention than the tokens ever will.
Sort tasks by what a wrong answer costs
The instinct is to sort by difficulty, sending hard-sounding work up and easy-sounding work down. That is the wrong axis, because difficulty is a property of the task and cost is a property of the mistake.
Put a task in the cheap lane when you were going to read the output anyway and a bad result is obvious within seconds. First-pass summaries of documents you know, tidying transcripts, extracting fields from a form, classifying an inbox, drafting something you will rewrite in your own voice regardless. The failure mode here is visible and recoverable, and you are the review step whether the model is expensive or not.
Put it in the mid lane when the output leaves your desk in something close to the form the model produced. Client-facing drafts, anything that goes into a document with your name on it, anything where a plausible-sounding error survives a quick skim. Terra at $2 and $12 [2] or Sonnet 5 at $2 and $10 [1] sit here, and the money buys fewer of the small errors that are expensive precisely because they look fine.
Reserve the top lane for work where a mistake is hard to reverse or hard to see. Contract review, reasoning chains where an early wrong step contaminates everything after it, code that will run unattended. That is the twenty per cent that justifies Fable at $10 and $50 [1] or Sol at $5 and $30 [2], and protecting the budget for it is the actual argument for demoting everything else.
Two levers make the ladder cheaper without changing lanes. OpenAI prices cached input at $0.50 per million for Sol and $0.02 per million for Luna, against $5.00 and $0.20 uncached, and offers batch processing at 50% off inputs and outputs [2]. If your work has a fixed preamble that repeats across runs, or a queue that does not need an answer this second, those two together move more money than most model swaps do.
Open weights are the floor under the ladder
The open-weight lane is worth understanding even if you never use it, because it sets a ceiling on how badly the paid lanes can treat you.
The gpt-oss models are released under the Apache 2.0 licence as weights you download and hold, and the smaller of the two is specified to run on an edge device with 16 GB of memory [8]. Anything with those properties cannot be withdrawn from you, cannot be repriced, cannot have its data policy amended after you built on it, and cannot be retired on a vendor’s schedule. For a specific class of work, mainly bulk processing of material you would rather not send anywhere at all, that is the whole argument.
It is not free in the way the word suggests. You pay in hardware, setup time, and the ongoing job of keeping the thing running, and a model you host yourself has no support line. The honest framing is that open weights are cheap per unit and expensive per hour, which makes them a good fit for high-volume repetitive work and a poor fit for the occasional hard question. Most small operators should know the option exists, price it once against their actual volume, and in many cases correctly decide not to bother yet.
Make the lane a setting rather than a rewrite
All of the above is only worth anything if changing your mind is cheap, and that is a matter of how the work is stored rather than which model you picked.
Write down each recurring task with the lane it runs on and one line saying why. That document is the thing you actually maintain. When a free window closes or a price changes, you are editing a list rather than reconstructing your own decisions from memory. Keep prompts and reference material in files you own instead of inside one product’s saved projects, because a prompt you cannot paste into a second tool is a prompt you cannot price anywhere else.
Then keep a fixed test set. Twenty real inputs from your own work, with the outputs you were happy with, stored in a folder. When a lane changes, you run those twenty through the candidate and read the results next to the old ones. This is not a benchmark and it does not need to be. It exists so that “the cheap model is fine for this” is a claim you have checked against your own material rather than against a chart someone published.
Set a review date rather than a monitoring habit. Once a quarter, open the pricing pages of the vendors you use, check whether any of your lanes moved, and re-run the test set on the two tasks with the highest volume. Fifteen minutes, four times a year, catches almost everything that matters.
What still goes wrong
The clean multiples in the price tables make routing look more precise than it is. A model that is a fifth of the price is not a fifth as good at your particular task, and it is not four fifths as good either. The relationship is lumpy and specific to the work, which is why the twenty-input test set matters more than any published comparison. Nothing here tells you where the line falls for your material, only how to find it cheaply and how to check it again later.
The bigger risk is the slow drift you do not notice. A demoted task keeps producing output that looks right, and the errors are small enough that no single one is worth mentioning, and three months later the summaries have quietly stopped catching the caveat that used to matter. Reading two outputs side by side once is not evidence that a lane is safe. If the work carries real consequences, keep a sample from before the switch and compare against it after a month of ordinary use, rather than trusting a first impression formed on the day you were motivated to save money.
Finally, free lanes end without warning and free is not the same as unmetered. Free routes carry request caps [7], free API tiers carry data-use terms that make them unsuitable for a large share of professional work [4], and free tiers on consumer products are the vendor’s most adjustable line item. None of that makes them a bad deal. It makes them a discount rather than a foundation, and the only real defence is the one this guide is about: keeping the cost of moving a task between lanes low enough that being wrong about any single lane stops mattering very much.
- 01Anthropic — Claude pricingclaude.com
- 02OpenAI — API pricingopenai.com
- 03Google — Gemini API pricingai.google.dev
- 04Google — Gemini API Additional Terms of Serviceai.google.dev
- 05OpenAI Help Center — How your data is used to improve model performancehelp.openai.com
- 06Anthropic Privacy Center — Is my data used for model training?privacy.claude.com
- 07OpenRouter — API rate limitsopenrouter.ai
- 08OpenAI — Introducing gpt-ossopenai.com