How to read an AI model announcement
Sort every AI announcement into four kinds, so you can ignore user counts and roadmap teasers and act only on shipped models, price changes and retirement dates.
on this page · 0 / 0 checked
An AI vendor publishes something most weeks, and the post is written to make you feel behind. A user-count milestone. A model with a decimal bump in its name. One sentence about a training run for a model that does not exist yet. You read it on a Tuesday, wonder whether the setup you built in March is now the wrong one, and lose an hour to a question the announcement was never capable of answering.
Almost none of it requires you to do anything. A small number of announcements do, and those are rarely the ones that get forwarded to you. What follows is a sorting rule: four kinds of announcement, what each can and cannot tell you, and the two kinds that arrive with a date you have to meet. It is written for someone running a business on tools they pay for. If you are building a product on a frontier lab’s roadmap under NDA, you already have better information than any blog post and a different set of problems.
Four kinds of announcement, and only two have deadlines
Every AI announcement is one of four things, and you can tell which within about ten seconds of reading it.
Distribution news is about the vendor: user counts, revenue, funding, a partnership, a chip deal. Roadmap news is about a model that does not exist yet: a training run, a teaser, a name with no product behind it. Shipped news is a model you can call today, which means it has an exact model ID and a price on a pricing page. Lifecycle news is about something you already use changing: a price rising, a version being retired, a policy shifting.
The two questions that sort them are boring and they work. Can I call this today, and what does it cost. Is something I already call going away or changing price. If an announcement answers neither, it is context, not a task. Read it, form an opinion if you like, and change nothing.
User counts tell you about the vendor’s staying power, not your workload
In his remarks on Alphabet’s Q2 2026 results, published on 22 July 2026, Sundar Pichai referred to “the Gemini app, which now has 950 million monthly active users” [8]. That is a real number and a large one, and it is the kind of number that travels furthest, because it is legible without any technical context.
Be precise about what it measures, and about what it covers. Monthly active users is a distribution and integration metric, and the same remarks show why scope matters: AI Mode in Search is reported as a separate line, which Pichai put past 1 billion monthly active users since its global expansion the previous October [8]. Two surfaces, two headline figures, and neither one measures how good any Gemini model is at the specific job you would give it. Neither tells you which model tier those people are reaching, either.
There is one decision distribution news legitimately feeds, and it is worth naming because dismissing this category entirely is also wrong. If you are about to build something you cannot easily move, a vendor’s scale is evidence about whether the vendor will still be there in three years and still supporting the thing you built. That is a durability judgement, and it is a fair use of a user count, a funding round or a revenue figure. It is not a quality judgement. Keep the two apart, because the same headline gets used to argue both and only one of them holds.
A roadmap teaser is a line in a file, not a reason to wait
On 21 July 2026, Google published a post announcing Gemini 3.6 Flash, Gemini 3.5 Flash-Lite and Gemini 3.5 Flash Cyber. Inside it, one sentence: “We have started our most ambitious pre-training run yet, for Gemini 4” [4]. The remainder of that sentence says only that Google is pleased with the progress. No date, no price, no capability claim, no access tier.
By early September, here is what a developer can actually call. Google’s model list carries gemini-3.8-flash as a stable model, described by Google as its most intelligent Flash model, engineered for long-horizon software engineering, autonomous agents and complex enterprise workflows [2]. The top of the Gemini 3 line, gemini-3.1-pro-preview, is still labelled preview [2]. The older gemini-3-pro-preview has moved to the previous-models list and is marked shut down [2], with a shutdown date of 9 March 2026 on the deprecations page [3]. Gemini 4 has a name and a training run and nothing you can put in a request.
The durable rule is short. You cannot benchmark a model you cannot call. An unreleased model is not a reason to defer a decision you need to make this month, because there is nothing to compare it against and no date on which there will be. It is a reason to keep the decision reversible, which for most of this work means avoiding contracts and architectures that assume one specific model forever. It is also worth noticing where that sentence sat. A roadmap line dropped into a mid-tier release post holds attention while the actual shipping happens one tier down.
Shipped means it has a model ID and a price
The pricing page and the model list are the ground truth, not the blog post. The blog post is written to be quoted; the pricing page is written by people who have to bill you correctly.
Two things become obvious once you read them. First, the spread inside a single vendor’s own lineup is larger than the spread between vendors. On Google, gemini-3.5-flash-lite is $0.30 per 1M input tokens and $2.50 per 1M output, while gemini-2.5-pro is $1.25 per 1M input for prompts up to 200,000 tokens, $2.50 above that, and $10.00 to $15.00 output across the same threshold [1]. On OpenAI, at standard short-context rates, gpt-6-astra is $10.00 input and $50.00 output per 1M tokens, and gpt-5.6-luna is $0.20 and $1.20 [7]. Astra costs 50 times as much per input token as Luna, inside one company’s catalogue. Which model you route to is a much bigger lever on your bill than which company you buy from, and no announcement will ever tell you that, because no vendor wants to point at its own cheap tier.
Second, the version you name has consequences. Google documents four kinds of version string. Stable versions “usually don’t change”, and the documentation says most production apps should use a specific stable model [2]. Preview versions may be used in production and typically have billing enabled, but “will be deprecated with at least 2 weeks notice” [2]. A latest alias such as gemini-flash-latest “will get hot-swapped with every new release of a specific model variation”, and for breaking changes a 2-week notice is provided through email before the version behind the alias changes [2]. Experimental versions are described as typically not suitable for production use, with more restrictive rate limits [2]. If you pinned an alias because it was shorter to type, you have quietly accepted a two-week window in which your outputs can change shape.
Promotional prices expire, and the expiry is in the footnote
The number you benchmarked on is often an introductory number, and the date it ends is already published, in small type, next to the price.
Gemini 3.8 Flash, 3.7 Flash and 3.6 Flash are currently listed at $0.75 per 1M input tokens through 31 December 2026, rising to $1.50 on 1 January 2027, with output moving from $3.75 to $7.50 on the same date [1]. That is a doubling with four months of notice sitting in plain sight on the pricing page. OpenAI does the same thing in the other direction on its own page, noting that GPT-5.6 Sol’s promotional pricing is available at least through 21 November 2026 [7]. Neither of these was a headline. Neither will be.
Three other line items on those pages move real money and never appear in announcements either. Batch processing runs at half the standard rate at both vendors, for work that does not need an answer in the next few seconds [1][7]. Cached input is far cheaper than fresh input, with gpt-6-astra cached at $1.00 against $10.00 standard [7]. And options you might assume are free are not: OpenAI charges a 10% uplift on regional processing endpoints for eligible models released on or after 5 March 2026 [7], and Google’s long-context tier on gemini-2.5-pro doubles the input rate above 200,000 tokens [1].
input millions × input price, plus output millions × output price. Run it again at the post-promotion price to see what the same volume costs after the expiry date. Computed in the page; nothing is sent anywhere.
Deprecation notices are the only announcements with a clock
This is the category that actually generates work, and it is the one nobody forwards. Each of the three major vendors publishes a deprecation page, states a notice period, and lists retirement dates [3][5][6]. Read those three things once per vendor and you will know how much warning you are entitled to.
Anthropic notifies customers with active deployments and commits to “at least 60 days’ notice before model retirement for publicly released models”, running models through four stages: active, legacy, deprecated and retired [5]. Its published table shows the policy being met rather than exceeded, with claude-opus-4-1-20250805 deprecated on 5 June 2026 against a retirement date of 5 August 2026 [5]. OpenAI publishes longer minimums: at least 6 months for generally available models, at least 3 months for specialized variants, and preview models that “may be retired with much shorter notice, such as 2 weeks”, with a stated exception where safety or compliance requires a faster retirement [6]. Its June 2026 entry notified developers on 11 June 2026 that older GPT-5 and o3 snapshots would be removed from the API on 11 December 2026, naming gpt-5.6-sol, gpt-5.6-terra and gpt-5.6-luna among the replacements [6]. Google takes a third approach, publishing shutdown dates it describes as “the earliest possible dates on which a model might be retired”, with the exact date communicated to users in advance [3].
The practical consequence is one page you maintain and one subscription you set up. The page lists every exact model ID your business calls and where each call lives, including the ones inside automations you built and forgot. The subscription is each vendor’s deprecation page, checked monthly or watched however you watch pages. When a notice lands, you do not have an opinion to form, you have a dated task: swap the ID, run your own examples through the replacement, confirm the output still passes, and do it before the date rather than on it. Sixty days sounds generous until the model ID turns out to be buried in an automation nobody remembers building.
What still goes wrong
Every price and policy quoted here is a snapshot verified on the date in the header, which is the point rather than a caveat. Tiers get renamed, promotional windows get extended or quietly dropped, and notice periods are policy statements a vendor can revise. Follow the links before you commit money, and treat any figure quoted in a blog post, including this one, as the starting point of a check rather than the end of it.
The sorting rule also cannot tell you the thing you most want to know, which is whether a newly shipped model is better at your work. Benchmarks will not tell you either, since they measure tasks that are not yours. The only cheap answer is a folder of 10 to 20 real tasks with outputs you already judged good, saved somewhere you can find them, run through the new model in half an hour. Most people who worry about model announcements do not have that folder, and building it is worth more than any amount of announcement-reading.
There is one more gap worth being honest about. Every guarantee quoted in this guide comes from API documentation: the version strings, the notice periods, the published retirement dates [2][5][6]. Those commitments are written for developers calling a model by its exact ID. If the precise shape of an output matters to your business, that is an argument for doing the work against a pinnable API model rather than inside a chat interface where you cannot name a version. And this whole approach is deliberately biased toward ignoring the future, which is correct when you are spending your own money on tools that must work this quarter. It is the wrong bias if you are in a market where being first on a new capability is the entire advantage. Know which of those you are before you decide that a roadmap teaser is safe to skip.
- 01Google — Gemini API pricingai.google.dev
- 02Google — Gemini API models and version namingai.google.dev
- 03Google — Gemini API deprecationsai.google.dev
- 04Google — Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyberblog.google
- 05Anthropic — Model deprecationsplatform.claude.com
- 06OpenAI — Deprecationsdevelopers.openai.com
- 07OpenAI — API pricingdevelopers.openai.com
- 08Sundar Pichai — Alphabet Q2 2026 earnings, CEO remarksblog.google