tuesday, october 6, 2026 · the day's ai, attributed published by trilot llc · wyoming
guide · running the business

The model you built on has a retirement date

How to find every model string you are pinned to, work out what staying put costs, and migrate before the vendor's retirement date arrives.

Published 2026-09-05 · Updated 2026-09-05 · Read 9 min · Reviewed by Rami Steitieh

Verified 2026-09-05 · Rami
on this page · 0 / 0 checked

Anthropic retired claude-opus-4-1-20250805 on August 5, 2026. It had notified developers on June 5, 2026, 61 days earlier, one day past its own 60-day floor [2]. Claude Sonnet 4 and Claude Opus 4 went dark on June 15, 2026 [2], and Google shut down gemini-2.0-flash on June 1, 2026 [7]. If one of those strings was sitting in your code, in an automation step, or in a config file you have not opened since last year, you found out when the requests started failing.

The usual mental model is that a superseded model quietly gets cheaper and you can sit on it as long as you like. That is not what happens. Old models mostly hold the price they launched at until the day they are switched off, and the price cut arrives attached to a new model name. This guide is for you if there is a model string somewhere in your business that you own: in code, in a Zapier, Make or n8n step, in a Cursor setting, in a notebook. If you only use Claude, ChatGPT or Gemini through the chat window, skip it. Your vendor migrates you, and the worst thing that happens is that the answers change.

The new model is the cheap one

Claude Sonnet 5 costs $2 per million input tokens and $10 per million output tokens. The two Sonnet versions before it, Claude Sonnet 4.6 and Claude Sonnet 4.5, still cost $3 and $15 [1]. The older model is 50% more expensive per token for the same job.

Google is the same shape and steeper. Gemini 3.8 Flash, Gemini 3.7 Flash and Gemini 3.6 Flash all cost $0.75 per million input tokens and $3.75 per million output. Gemini 3.5 Flash, one rung older, is $1.50 and $9.00 [5]. Staying on 3.5 Flash costs double on input and 2.4 times on output.

At the top of Anthropic’s range the gap is wider. Claude Opus 4.5, 4.6, 4.7, 4.8 and Opus 5 are all $5 and $25. Claude Opus 4.1 and Claude Opus 4, the generation before them, were $15 and $75, and both are retired on the Claude API [1][2]. The top of the range got 3 times cheaper per token. It did that by shipping under a new name, not by discounting the old one.

This is the part that catches people who did everything else right. You pinned a model for stability, you watched the market, you saw the headlines about prices falling, and you assumed the falling price would reach you. It does not reach you. It sits one model string away, and the only way to collect it is to move.

Pinning is correct, and it has a bill and a deadline

None of this means aliases are safer. Google’s own documentation tells you to use specific stable versions in production rather than a scheme that auto-updates, and warns that you get 2 weeks of email notice before the model behind a latest alias changes underneath you [6]. A pinned string is still the right call for anything a customer touches.

What the pin costs you is that nothing improves and nothing gets cheaper until you act, while a clock you did not set is running. Anthropic’s deprecation table currently records Claude Opus 4.1 retired on August 5, 2026, Claude Sonnet 4 and Claude Opus 4 on June 15, 2026, Claude 3.7 Sonnet on February 19, 2026, and both Claude 3.5 Sonnet snapshots on October 28, 2025 [2]. OpenAI has gpt-3.5-turbo-0125, gpt-4-0613 and o1 shutting down on October 23, 2026, and its GPT-5 and o3 snapshots on December 11, 2026 [4].

Treat a pinned string as a dated commitment rather than a decision you finished making. On the day you pin it, write the vendor’s retirement date next to it in whatever place you will actually look. Anthropic points you at your usage page to find which deprecated models you are still calling [2]; the same audit on any vendor takes an afternoon, and it is the cheapest afternoon in this guide. The calculator below turns the gap into a monthly number, which is usually what makes the migration get scheduled.

Your notice period is a vendor choice you already made

Anthropic commits to at least 60 days of notice before retiring a publicly released model [2]. OpenAI commits to at least 6 months for generally available models, at least 3 months for specialized variants such as its chat, Codex and deep research versions, and says preview models can go with as little as 2 weeks [4]. Google says preview models are deprecated with at least 2 weeks of notice [6].

That is a spread from 2 weeks to 6 months, and it is not a footnote. It is how much runway you get to notice the email, run your evals, fix whatever broke in the prompt, and ship. A 60-day window is comfortable if you already have a test set. It is not comfortable if the first task in the migration is writing one.

Preview models are where this bites hardest. Gemini 3.1 Pro and Gemini 3 Flash are both preview at the time of writing [6], and preview status carries the shortest notice at both Google and OpenAI [6][4]. Using a preview model to find out whether something is possible is sensible. Building a customer promise on one means you have accepted a 2-week eviction notice on that promise, whether or not you noticed signing for it.

Promotional prices come with a date printed on them

OpenAI’s pricing page states that GPT-5.6 Sol’s promotional pricing is available at least through November 21, 2026 [3]. Google’s states that Gemini 3.8 Flash’s $0.75 and $3.75 rate runs through December 31, 2026 [5]. Those are the numbers a spreadsheet built this week would use, and both of them expire inside the next 4 months.

If your margin per customer only works at a promotional rate, that is not a margin, it is a countdown. Put the expiry date in the same place as the retirement dates and run the spreadsheet again at the standard price to see what survives.

Not every price cut requires changing models. Anthropic’s Batch API is a 50% discount on input and output, which takes Claude Sonnet 5 to $1 and $5 [1], and Google’s paid tier lists its Batch API as a 50% cost reduction [5]. Prompt caching is the larger lever. On Anthropic a cache read costs 0.1 times the base input price, and 0.025 times on Claude Fable 5.1 and Claude Mythos 5.1 [1]. That is a 90% cut on every repeated input token, against the 33% you collect by moving from Claude Sonnet 4.6 to Claude Sonnet 5 [1]. If the same long brief or document goes into every call you make, cache it before you go shopping for a cheaper model.

Migration is not always a change to the model string

Anthropic has deprecated the temperature, top_p and top_k parameters on Claude Opus 4.7 and later, and setting them to non-default values returns a 400 error [2]. A request body that worked against every earlier Claude model now fails against the newer one. The fix is not a swap. It is rewriting the prompt to do the job the temperature setting was doing.

Whole platforms retire as well. OpenAI shut down its Assistants API on August 26, 2026, and its Evals platform, Agent Builder and reusable prompts API are scheduled to shut down on November 30, 2026 [4]. If your workflow lives inside one of those, the migration is a rebuild with a deadline attached, not a config change.

Context pricing is the third surprise. GPT-6 Astra costs $10 and $50 per million tokens on short context and $20 and $75 on long context [3]. Gemini 3.1 Pro is $2 and $12 up to 200,000 tokens and $4 and $18 above that [5]. Move to a model with a bigger window, start filling the window because you now can, and the bill moves in a direction your per-token comparison never predicted. Compare cost per finished task across the two models, on your real inputs, before you commit.

What you shipped is in somebody’s free plan

Claude’s Free plan includes Sonnet and Haiku, with Opus arriving on Pro and Max [8]. On the Gemini API, the free tier covers Gemini 3.8 Flash, 3.7 Flash, 3.6 Flash, 3.5 Flash, both Flash-Lite models and Gemini 2.5 Pro, among others, while Gemini 3.1 Pro Preview has no free tier at all [5].

Gemini 3.8 Flash became generally available on September 2, 2026 [7], at the same price as Gemini 3.7 Flash and half the price of Gemini 3.5 Flash [5]. That is the rhythm, and it does not appear to be slowing. Whatever capability you were charging for a year ago, assume a free user has a version of it in a chat window now, and check that your price is attached to something else.

The something else is the part that survives a migration: the prompt you tuned against real cases, the data you feed it, the checks you run on the output before it reaches anyone, the workflow it sits inside, and the customer who trusts you to have done all four. A model string is not a moat. It is a dependency with an end date, and the labs publish the date in advance.

The maintenance rhythm that keeps this cheap

Keep a small eval set. 20 to 40 real inputs with the output you consider correct is enough, and it converts migration from guesswork into an afternoon. Build it while nothing is on fire, because the week a retirement notice lands is the worst possible week to start.

Audit once a quarter, and audit everywhere, not just the repository. Model strings hide in Zapier and Make steps, n8n nodes, Cursor settings, scheduled scripts, and the one internal tool nobody owns. For each string, record the current price, the price of the model that replaced it, and the retirement date.

Run a mixed stack rather than picking one model for everything. Claude Haiku 4.5 is $1 and $5, and Claude Fable 5.1 is $10 and $50 [1]. GPT-5.6 Luna is $0.20 and $1.20, and GPT-6 Astra is $10 and $50 [3]. That is a 50-fold range on input price inside a single vendor. Classification, extraction and routing rarely need the top rung. Reserve it for the tasks where you can show the cheaper model failing.

Track cost per finished unit of work, not the monthly total. A cheaper model that needs two attempts, or produces output someone has to correct, is more expensive than its token price. The monthly bill hides that. Cost per completed task does not.

checklist
Quarterly model audit
0 of 8 · saved in this browser only
calculator
What staying on the old model costs
— $ / month

Defaults compare Claude Sonnet 4.6 at $3 and $15 against Claude Sonnet 5 at $2 and $10 [1]. Computed in the page; nothing is sent anywhere.

What still goes wrong

Every number here was read off a vendor page on September 5, 2026, and two of those pages carry their own expiry dates [3][5]. Prices, model names and retirement dates move faster than any guide can. Use the structure and re-read the pricing page; do not quote this article’s figures back at your accountant in six months.

Cheaper is not reliably cheaper. A smaller model that retries, or that produces output a human has to repair, can cost more in total than the model you left. The token price is the only part of that which is published, and it is the part least likely to decide the answer. Measuring cost per finished task on your own inputs is the only thing that settles it, and it takes real work to set up.

The published notice period is a floor, not a promise about capability. Anthropic defines a legacy model as one that no longer receives updates and may be deprecated later [2], which means “still responds” and “still supported” are separate states, and you can be in the first without the second for a long time. Nothing in any vendor’s policy obliges them to keep a specific behaviour you depend on, only to warn you before the endpoint stops answering. If that risk is genuinely unacceptable for your business, the alternative is running open-weight models on hardware you control, which trades a retirement date for an operations burden. That is a real option and a much larger commitment than a quarterly audit, and this guide does not cover it.

sources
  1. 01Anthropic — Claude model pricingplatform.claude.com
  2. 02Anthropic — Model deprecationsplatform.claude.com
  3. 03OpenAI — API pricingdevelopers.openai.com
  4. 04OpenAI — Deprecationsdevelopers.openai.com
  5. 05Google — Gemini API pricingai.google.dev
  6. 06Google — Gemini API models and version supportai.google.dev
  7. 07Google — Gemini API changelogai.google.dev
  8. 08Anthropic — Claude plans and pricingclaude.com
next guide
When a new rule kills a feature you already shipped
9 min · verified 2026-09-05
related guides