Frontier prices keep falling. That is not the same as your bill falling
See how model prices actually move, why a price hold changes your routing rather than your budget, and which jobs to re-test when it happens.
on this page · 0 / 0 checked
A new top-tier model lands and the announcement says the price has not moved. You read that as good news, because it is. Then nothing about your week changes. You keep sending the same work to the same model you picked a year and a half ago, on a spreadsheet you no longer have, and the announcement joins the pile of things you will get to.
The price hold is real, it repeats, and it is worth understanding. It also does not lower your bill by a cent, and it is not an instruction to upgrade anything. What it does is quietly invalidate one decision you made on cost grounds and never went back to. This guide is for someone who already pays for model access, by API or by subscription, and has at least two models available to send work to. It is not for someone with a procurement team negotiating committed-use discounts, and it is not for someone who has never run the same job through two models and compared the results.
The published record shows prices going down, then sitting still
Anthropic’s price list is the clearest place to watch the pattern, because it still shows the old rungs. Claude Opus 4 and Opus 4.1, both now marked retired, were listed at $15 per million input tokens and $75 per million output tokens [1][3]. Every Opus released since sits at $5 and $25: version 4.5, 4.6, 4.7, 4.8 and 5, five consecutive releases at one number [1]. When Opus 5 shipped on 24 July 2026, Anthropic’s own announcement spelled it out, describing the model as “priced at $5 per million input tokens and $25 per million output tokens (the same as Opus 4.8)” [2].
The mid tier moved the same way. Sonnet 4.5 and 4.6 are listed at $3 and $15, and Sonnet 5 at $2 and $10. Haiku 4.5 is $1 and $5. Above Opus, Fable 5, Fable 5.1, Mythos 5 and Mythos 5.1 all sit at $10 and $50 [1]. OpenAI’s list has the same shape and roughly the same rungs: gpt-6-astra at $10 and $50, gpt-5.6-sol at $4 and $20, gpt-5.6-terra at $2 and $12, and gpt-5.6-luna at $0.20 and $1.20 [7].
Read those two lists together and the useful structure appears. There are four or five price rungs, each roughly a factor of two to five apart, and a given rung tends to hold its number while the model occupying it gets replaced. That is the actual mechanism behind a price hold. You are not being given a discount. You are being given a better tenant in the same apartment, and the rent notice is the announcement telling you the tenant changed.
Falling is a tendency, not a rule, and the page has footnotes
It is tempting to turn the last two paragraphs into a law and plan around it. The price pages themselves say otherwise, in writing, on the same screens.
Google’s Gemini API pricing lists Gemini 3.8 Flash at “$0.75 through December 31, 2026” for input and “$1.50 starting January 1, 2027”, with output going from $3.75 to $7.50 on the same date and context caching from $0.075 to $0.15 [8]. That is a published doubling with a date attached. OpenAI’s pricing page carries a similar note in the other direction, saying that “GPT-5.6 Sol’s promotional pricing is available at least through November 21, 2026” [7], which makes $4 and $20 a floor with an expiry rather than a standing rate.
Shape matters as much as level. Gemini 3.1 Pro Preview is priced at $2 input and $12 output for prompts at or under 200,000 tokens, and $4 and $18 for prompts above that [8]. Same model, two prices, decided by how much you paste in. OpenAI tiers the same way, listing gpt-6-astra at $10 and $50 at standard length and $20 and $75 on long context [7]. Anthropic goes the other way and includes the full 1M-token context window at standard pricing on Claude 4.6 and later, with no per-token increase for larger contexts [1]. Before you build a forecast on any of these, read the conditions under the number, and note the date.
Price per token is the smallest lever you have
The rate is the part you cannot negotiate and the part you spend the most time reading. The two levers you can pull are both larger.
Prompt caching is the first. Cache read tokens cost 0.1 times the base input price on every model except Fable 5.1 and Mythos 5.1, which read at 0.025x, so Opus 5 input drops from $5 to $0.50 per million tokens on cached content, and Sonnet 5 from $2 to $0.20 [1][4]. Writing to the cache costs 1.25x base input for the 5-minute cache and 2x for the 1-hour version [1][4]. The detail that catches people is the minimum: cacheable prompts must clear a per-model threshold, 512 tokens on Opus 5, 1,024 on Opus 4.8 and Sonnet 5, 2,048 on Opus 4.7, and 4,096 on Opus 4.6, Opus 4.5 and Haiku 4.5. A prompt below the line is simply processed without caching, and the documentation is explicit that “no error is returned” [4]. You lose the 90 percent discount silently.
The Message Batches API is the second. All batch usage “is charged at 50% of the standard API prices”, most batches complete within an hour, and a batch expires if processing has not completed within 24 hours [5]. That puts Opus 5 at $2.50 and $12.50 in batch, and Sonnet 5 at $1 and $5 [1][5].
Now put those next to each other. Opus 5 input runs from $5 at list, to $2.50 in batch, to $0.50 on a cache read [1][4][5]. The distance between careless and careful use of a single model is wider than the distance from Opus 5 down to Sonnet 5 on the same list. Anyone comparing two vendors’ headline rates while sending uncached, unbatched, real-time requests is optimising the wrong variable.
The largest lever is not on any price page. It is how many tokens you send. A rate that has held across five releases buys you nothing if your usage moved from one prompt with a paragraph in it to an agent loop that reads twelve files and retries twice. Bills rise while rates fall, and the reason is almost always volume.
What a price hold actually changes is your routing
Somewhere in your setup is a job running on a cheaper model for a reason you can no longer reconstruct. It was probably cost, decided against a price that no longer exists. That is the one thing a price hold makes stale, and it is the only thing worth acting on.
The re-test is narrow and mechanical. Pick one job. Pull 20 real inputs from the last month, not invented ones, because synthetic examples are always cleaner than what actually arrives. Decide in advance what better means for that job, in a sentence, before you see any output. Run all 20 through both models. Grade them without knowing which is which if you can manage it. Then price the winner.
The Sonnet 5 to Opus 5 gap on Anthropic's list price: $3 more per million input tokens, $15 more per million output. Computed in the page; nothing is sent anywhere.
Run it with your own volumes before you argue about it. For a job at a few hundred runs a month with ordinary prompt sizes, the answer is usually tens of dollars, which is less than an hour of your time. At that size the decision is settled by the 20 cases and not by the arithmetic. If you cannot see a difference across 20 real inputs, the cheaper model was the right call and the announcement was not about you.
If you pay a subscription, none of this appears on your invoice
Most people at this size never touch the meter. Anthropic’s plan page lists Free at $0 with Sonnet and Haiku and a 200k context window; Pro at $17 a month on an annual subscription or $20 billed monthly, which adds Opus access, Projects and Claude Code, and gives “at least 5x more usage per 5-hour session than Free”; Max from $100 a month for “5x or 20x more usage per 5-hour session than Pro”; and Team seats at $20 a month annual or $25 monthly for standard, $100 or $125 for premium. Every plan resets on a rolling five-hour session window, and paid plans add weekly limits on top [6].
Nothing in that list moves when a per-token price holds. What a cheaper-to-serve top model does to a subscription is change how much of it fits inside a fixed cap, which is a genuine benefit and an entirely unpublished one. You cannot budget against it and you cannot verify it from outside.
The practical consequence is that subscription users and API users should read the same announcement differently. On a subscription, the question is never whether the model is worth the money, because the money is fixed. It is whether you are hitting the weekly cap, and if you are, whether the answer is the next tier up or moving the heavy repetitive work onto the API, where it is metered, visible, and eligible for the batch and caching discounts that a chat window cannot give you [1][5].
The date that forces your hand is the retirement date
Prices are optional to act on. Retirements are not, and they run on a published schedule that most people never open.
Anthropic states that it provides “at least 60 days’ notice before model retirement for publicly released models”, and it does retire them. claude-opus-4-1-20250805 was deprecated on 5 June 2026 and retired on 5 August 2026, with claude-opus-4-8 named as the replacement. claude-opus-4-20250514 and claude-sonnet-4-20250514 were both deprecated on 14 April 2026 and retired on 15 June 2026. Requests to models past the retirement date fail [3].
The same table publishes tentative dates for models in service today: claude-sonnet-4-5-20250929 not sooner than 29 September 2026, claude-haiku-4-5-20251001 not sooner than 15 October 2026, claude-opus-4-5-20251101 not sooner than 24 November 2026, and claude-opus-5 not sooner than 24 July 2027 [3]. If you pinned an exact model string in a script, a Zap or a prompt library, you accepted one of those dates without noticing.
So the calendar to work from is the retirement table, not the price list. A price change invites a decision. A retirement schedules one for you, and it arrives whether or not you have 20 test cases ready. Keeping every model string in one file you can find turns that from an emergency into an afternoon.
What still goes wrong
The re-test is the step that does not happen. Everyone agrees it is sensible, it costs an afternoon, and it might return a result you would rather not have. Meanwhile two years of routing decisions sit in a stack, each one made against a price that has since changed, and no single one of them is worth an afternoon on its own. The only version of this that survives contact with a real week is a small standing habit: one job, twice a year, 20 real inputs. Not a review of the whole stack.
Published rates are also not the whole cost of a switch. A prompt tuned against one model does not always carry over, and a model that follows instructions more literally can break a workflow built around its predecessor’s habits. Vendor benchmarks measure what the vendor chose to measure, so “better at the same price” is a claim about an average across an eval suite you did not design, and your work is not that average. The 20 cases are the only evidence that settles it for you.
And a price hold is a vendor’s decision about margin and market share, not a property of the technology. Every list price quoted here is a snapshot of pages published today and every one of them can move. Two of them already have movement scheduled in writing, in Gemini 3.8 Flash’s step up on 1 January 2027 [8] and the “at least through November 21, 2026” wording on OpenAI’s Sol pricing [7]. Check the page before you quote the number, including the numbers in this guide.
- 01Anthropic — Claude model pricingplatform.claude.com
- 02Anthropic — Introducing Claude Opus 5anthropic.com
- 03Anthropic — Model deprecationsplatform.claude.com
- 04Anthropic — Prompt cachingplatform.claude.com
- 05Anthropic — Message Batches API (batch processing)platform.claude.com
- 06Anthropic — Claude plans and pricingclaude.com
- 07OpenAI — API pricingdevelopers.openai.com
- 08Google — Gemini API pricingai.google.dev