Why cheaper models don't make a cheaper bill
Find out where your AI spend actually goes, take the three discounts vendors already offer, and stop paying full price for the same prompt twice.
on this page · 0 / 0 checked
Every few months a vendor cuts the price of a model and the number is genuinely large. Then you open last month’s invoice and it is the same as the month before, or higher. Nothing is wrong with the announcement. The cut is real, the number on the pricing page is real, and the savings went somewhere. They mostly did not go to you, because the thing that got cheaper is not quite the thing you were buying.
This guide is about closing that gap. It covers where the money actually goes when you run AI through an API or an automation platform, which of the vendors’ discounts you have to go and claim yourself, and how to read the next price cut so you know whether to expect another one. It is written for someone who sees a variable AI bill and cannot fully explain it. If your entire AI spend is one flat subscription, most of this does not apply, and the two sentences that do are at the end of the first section. Prices below are US list prices fetched on 4 September 2026, and vendors change them without much notice.
A token got cheaper, and your answer might not have
The published rates are low and they are still falling. On the OpenAI API, GPT-5.6 Sol is $5.00 per million input tokens and $30.00 per million output; Terra is $2.00 and $12.00; Luna is $0.20 and $1.20 [1]. On the Claude API, Fable 5.1 is $10 and $50 per million, Opus 5 is $5 and $25, Sonnet 5 is $2 and $10, Haiku 4.5 is $1 and $5 [2]. Gemini 3.8 Flash is $0.75 and $3.75 [3].
Notice which side is expensive. Output costs six times input on Sol, five times on Fable 5.1 and Sonnet 5 [1][2]. That matters more than it used to, because reasoning models spend tokens thinking before they answer, and those tokens are billed as output tokens even though you never see them [8]. So the price per million tokens is not the price of your answer. A model that costs half as much per token and reasons for twice as long costs you exactly the same, and the pricing page will still be advertising a 50% cut.
The number that moves your bill is cost per finished task: one support ticket classified, one document summarised, one pull request reviewed. Work it out once for your three most common jobs, from a real usage response rather than an estimate, and you will have something you can compare across models and across months. Per-million rates are for comparing vendors. Cost per task is for comparing decisions.
If you are on a flat subscription instead, the arithmetic is different and shorter. A price cut reaches you as more capability at the same price, not a refund. Claude Pro is $20 per month billed monthly, or $17 per month on the annual plan at $200 up front [2]. Your only levers are which tier you sit on and how many subscriptions you are quietly paying for.
Three things get cheaper, and only one of them is yours
When OpenAI explained where the GPT-5.6 savings came from, it named three separate layers: the models themselves, which “take a more direct path through work”; “the inference systems that run them”; and “the agentic harness that connects them to tools and context” [7]. It gave figures for the middle layer, where kernel work “helped reduce the end-to-end cost of serving the model by 20%” and other experiments “increased token-generation efficiency by more than 15%” [7]. The third layer it described as “smarter context management” that helps “agents avoid repeating completed work” [7].
That split is the useful part, and it survives whichever vendor announces next. Savings from the first two layers arrive in the price you are quoted, automatically, whether or not you change a line of code. You do nothing and the rate falls. Savings from the third layer arrive only if the harness belongs to the vendor. When you use the Claude or ChatGPT app, or a vendor’s own coding agent, their context management is doing that work for you. When you built the loop yourself, in n8n or in your own code, that layer is yours, and no announcement will fix an agent that re-sends the same 30,000-token brief on every one of its twelve steps.
The split also gives you the question to ask of the next announcement. Which layer moved, and can that same layer move again. A rate cut that came out of serving costs shows up on the pricing page you already read. A cut that came out of a redesigned harness only reaches the customers who use that harness, and it says nothing about the one you wrote.
Caching is the discount you have to earn
Every serious API now charges roughly a tenth for input it has seen before, and it is easy to lose that discount without ever seeing an error.
On the OpenAI API, caching is on by default: “Prompt caching is enabled by default for supported OpenAI models” [4]. The minimum is 1,024 visible input tokens for GPT-5.6 and later, cached input costs 0.1 times the uncached rate, and a cache write costs 1.25 times, so a prefix written once and reused once costs 1.35 times its ordinary input cost “compared with 2× for processing it twice” [4]. A cached prefix stays eligible for at least 30 minutes after its most recent write or reuse on GPT-5.6 and later [4]. Anthropic makes you opt in, either by adding one cache_control field at the top level of the request, which applies the breakpoint to the last cacheable block and moves it forward as the conversation grows, or by placing cache_control on individual blocks yourself [5]. Its rates: a 5-minute cache write is 1.25 times the base input price, a 1-hour write is 2 times, reads are 0.1 times on most models and 0.025 times on Fable 5.1 and Mythos 5.1, and the minimum cacheable prefix runs from 512 tokens on the larger models to 4,096 on Haiku 4.5 [5]. Gemini charges $0.075 per million for cached reads on 3.8 Flash against $0.75 standard, plus $0.50 per million tokens per hour of storage [3].
The way people lose it is always the same. Reuse requires the entire rendered prefix to match, so a different model, a changed tool definition or its ordering, an edited schema, a different output format setting or a compaction step that rewrites earlier context all break the hit [4]. On Claude the invalidation cascades in order, tools, then system, then messages, so touching a tool definition throws away everything behind it [5]. The practical rule is one line long: stable content first, volatile content last. A timestamp, a session id or a customer name near the top of your system prompt quietly cancels the discount on everything after it, and nothing in your logs will say so.
So check the numbers rather than the intent. Both APIs report the cache in the usage object, as usage.input_tokens_details.cached_tokens on OpenAI [4] and cache_read_input_tokens on Claude [5]. If that count is near zero on a workflow that sends the same brief forty times a day, the cache is not engaging, and you are paying $2.00 per million where $0.20 was available [1].
Assumes cache reads at 0.1× the input rate and a warm cache, so it ignores the 1.25× write premium and any run that misses. 22 working days. Computed in the page; nothing is sent anywhere.
Half off for anything nobody is waiting on
The batch endpoints halve the rate, and the only thing they ask for is patience. OpenAI’s Batch API is a “50% cost discount compared to synchronous APIs”, with a 24-hour completion window, up to 50,000 requests per batch and a 200 MB input file limit [6]. Anthropic advertises the same 50% for batch processing [2]. Gemini’s batch rates are exactly half its standard ones, $0.375 input and $1.875 output on 3.8 Flash against $0.75 and $3.75 [3].
There is one condition, and it is not technical. Nobody can be waiting for the answer. That still leaves a large share of what most small operations actually run: classifying a month of support tickets, summarising a backlog, rewriting product descriptions, scoring a list of leads, backfilling tags on old content, generating the weekly digest that gets read at nine the next morning. Go and look at which of those are running synchronously in your own setup. Usually the reason is that the first version was a test someone ran by hand, and nobody went back to it.
The other half of the habit is worth breaking too. If a job runs on a schedule and its output lands in a table or a document, the schedule can submit the batch on one run and collect it on the next. The 24-hour window [6] is longer than most reporting jobs need it to be.
Route by difficulty, and look at the effort dial
The spread inside a single family is now wide enough to matter more than the spread between vendors. GPT-5.6 Sol’s output is $30.00 per million and Luna’s is $1.20, a factor of 25, from the same provider on the same day [1]. On Claude, Fable 5.1 output is $50 per million and Haiku 4.5 is $5 [2]. If everything in your system calls the top model because that is what you were testing with in week one, you are paying a premium on the large majority of calls that were never hard.
Newer can also be cheaper within a family, which cuts against the instinct to stay put, though you should read the dates. Gemini 3.8 Flash is $0.75 input and $3.75 output “through December 31, 2026”, then $1.50 and $7.50 “starting January 1, 2027”, while the older Gemini 3.5 Flash is $1.50 and $9.00 [3]. So today the newer model is half the price on input; in January it matches the older one on input and stays cheaper on output. Sitting on last year’s version is not the conservative choice it looks like, but part of the gap is a promotion with an expiry date printed on it.
Then there is the effort dial. OpenAI’s reasoning effort parameter accepts none, minimal, low, medium, high, xhigh and max, and “lower effort favors speed and lower token usage” [8]. Reasoning tokens are billed as output [8], which is the expensive side, so leaving effort high on a job that extracts a date from an email is a bill you wrote yourself. There is a hard stop available as well: max_output_tokens caps the total generated, including reasoning tokens [8]. Set it on anything running unattended.
What still goes wrong
Caching breaks invisibly. Nobody guarantees you a hit, and the list of things that void one is longer than it looks: changing the model, modifying tools or their schemas, ordering or descriptions, altering the structured output format, adjusting reasoning effort at request level, or running a compaction step that rewrites earlier context [4]. You can do everything correctly, upgrade a client library that reorders your request body, and lose the discount without a single error appearing anywhere. The only real defence is to treat a fall in the cached-token count as a bug and alert on it.
Cheap tiers do not announce their failures. A wrong answer from the cheap model arrives in the same shape and the same confident tone as a right one, and nothing in the response marks it. So routing by difficulty is only safe where you can check the result: a schema that has to validate, a number that has to reconcile, a person who reads it before it goes anywhere. For anything touching a customer, a contract or a money figure, treat the cheap tier as producing a draft rather than a result, and do the comparison on your own work rather than on someone’s benchmark, because the only quality that matters is quality on your inputs.
And none of this is a promise that the direction holds. The published rate is today’s rate, and it can be tomorrow’s higher rate on the same page: Google already lists Gemini 3.8 Flash input rising from $0.75 to $1.50 per million on 1 January 2027 [3]. Models get retired, tiers get renamed, and a discount that exists now can be restructured in a release note. If a workflow is only viable at a particular price, you do not have a workflow, you have a position. Keep the model name in one place in your configuration, price the job again when the model changes, and be honest that for a business whose whole AI spend is two flat subscriptions, none of these levers exists at all; there the only question is whether you need both.
- 01OpenAI — API Pricingopenai.com
- 02Claude — Pricingclaude.com
- 03Google — Gemini API pricingai.google.dev
- 04OpenAI — Prompt caching guidedevelopers.openai.com
- 05Anthropic — Prompt cachingplatform.claude.com
- 06OpenAI — Batch API guidedevelopers.openai.com
- 07OpenAI — Advancing the price-performance frontier with GPT-5.6openai.com
- 08OpenAI — Reasoning models guidedevelopers.openai.com