Why the model you are waiting for is late
Understand why a finished model can still be unavailable to you, and set your work up so a delay or a retirement costs you nothing.
on this page · 0 / 0 checked
A new model exists. The chart is circulating, someone has already written the thread about what it changes, and two of your clients have asked whether you are using it yet. Then you open the app and the thing you wanted is not there. Or it is there and the one capability you needed is missing, or sitting behind a tier you cannot buy, or open only to people who have cleared a check you have not heard of. The model is finished. You still cannot have the part of it you came for.
That gap used to be a rollout stagger, measured in days and mostly about capacity. It is now the visible end of a process, and the process runs in two places at once. Inside the lab, which grades its own model against a published capability framework and decides what to withhold from whom. And inside Washington, where a voluntary review window now sits between a lab and a broad release. This guide is for someone running a business of one to about ten people with real work sitting on a specific model. It is not for a lab, and not for a company that has a compliance function, which needs more than a reading of two policy pages. If nothing you own names a model version anywhere, nothing here can break for you and you can skip it.
A launch date is three dates now
OpenAI published the system card for GPT-6 Astra on 3 September 2026 [2]. Astra is, in OpenAI’s own words, “our first model to reach the Critical level of cybersecurity capability under our Preparedness Framework” [2]. OpenAI also calls it “the most capable model we have ever broadly deployed” [2], and broad is the operative word rather than complete. The card carries separate sections for “Trusted Access for Cyber” and “Trusted Access for Biology Research” [2], and lists internal controls as “stricter isolation, checkpoint encryption, universal monitoring of full trajectories including chains of thought (CoT), and a blocking alignment evaluation process” [2]. External evaluation was run by named third parties including the UK AI Security Institute, Apollo Research, SecureBio and Gray Swan [2].
Trusted access is the part that decides whether a given capability is yours. OpenAI describes Trusted Access for Cyber as “an identity and trust-based framework designed to help ensure enhanced cyber capabilities are being placed in the right hands” [7]. Users “can verify their identity at chatgpt.com/cyber” [7], and a further tier sits above that: “Security researchers and teams who may need access to even more cyber capable or permissive models to accelerate legitimate defensive work can express interest in our invite-only program” [7]. So the model ships broadly and some of what it can do does not.
Eight weeks earlier the same machinery produced a softer outcome. The GPT-5.6 system card, published 9 July 2026, covered “a new family of three models: Sol, our new flagship model; Terra, a capable lower-cost option; and Luna, our fastest and most cost-efficient model” [3], and determined that all three “warrant the same designations for our Preparedness Framework’s Tracked categories: High in Biological and Chemical, High in Cybersecurity, and below High in AI Self-Improvement” [3]. Fortune reported the same day that the family had been “initially limited to select partners following pressure from the Trump administration to stagger the release”, and quoted Sam Altman on government testing: “I think that’s good as long as the process is understandable, fair and quick” [8].
Three dates, then, and only one of them is yours. There is the announcement, which is a marketing event. There is restricted access, which goes to select partners [8] and to people who clear an identity or invitation gate [7]. And there is the date the capability you want is sold to somebody like you. Roadmaps written against the first date are fiction, and the honest version of a plan says which date it depends on.
The gate is a capability rating the lab gives itself
Nothing above was imposed by a regulator. The thresholds are the labs’ own, published in advance, and applied by the lab to its own model.
OpenAI’s are stated as capability levels, with Critical the level Astra reached on cybersecurity [2]. Anthropic runs the same shape under different names. Its Responsible Scaling Policy, version 3.4, effective 8 July 2026, states that “reaching certain Capability Thresholds requires us to upgrade our safeguards to the ASL-3 Security Standard or the ASL-3 Deployment Standard” [4], with the thresholds written around chemical, biological, radiological and nuclear weapons and around “the ability to fully automate entry-level AI research work” [4]. The policy describes its own approach to governance as “proportional, iterative, and exportable” [4].
The consequence for you is not really about frontier capability at all. It is that the thresholds are fixed and the models keep moving, so the population of gated models grows. OpenAI recorded the GPT-5.6 family as “the first time that smaller and faster members of a model family have received a High capability designation in any Tracked Category” [3]. The intuition that heavy safety machinery only ever applies to the expensive flagship, and that the budget model you route bulk work through will always be waved past, is an assumption with an expiry date on it.
Washington’s review is voluntary, and that is not the same as optional
Executive Order 14409, signed 2 June 2026, is the document to read rather than any summary of it. Section 3(a) directs the development of “a classified benchmarking process to assess the advanced cyber capabilities of AI models and determine the threshold at which an AI model should be designated a ‘covered frontier model’” [1]. Section 3(b) directs agencies to “design a voluntary framework with AI developers” under which developers would “provide the Federal Government with access to covered frontier models, subject to appropriate confidentiality, cybersecurity, insider-risk, and intellectual-property protection, use, and nondisclosure requirements, for a period of up to 30 days before they plan to release such models to other trusted partners” [1].
Read that clause slowly. The window runs before release to trusted partners, which is itself before anything reaches general availability. By the time a restricted preview appears, the government review may already have happened and finished.
Section 3(c) is equally worth quoting, because it is the line most coverage drops: “Nothing in this section shall be construed to authorize the creation of a mandatory governmental licensing, preclearance, or permitting requirement for the development, publication, release, or distribution of new AI models, including frontier models” [1]. There is no licence. There is no approval. A lab that shipped without participating would break no law.
None of this is a rule you comply with. It is a schedule you sit downstream of, and the distinction matters because it changes what you can do about it. You cannot file anything, appeal anything or qualify for anything. You can only stop building plans that assume a date.
What reaches you is a feature, not a model
Almost none of this arrives labelled. You are unlikely to be calling an API and reading version strings. You use Claude or ChatGPT or Gemini in a browser, Cursor in an editor, a Notion database, a Zapier step that quietly calls a model on your behalf. Capability gating reaches that stack as a mode that is missing, a limit that is lower than you remembered, or a feature that works for a company on an enterprise contract and not for you.
The vendors design for this openly. With GPT-5.6, OpenAI said “we provide an option in ChatGPT and Codex to easily retry prompts on lower-capability models” and described its deployment approach as “starting conservatively and improving based on what we learn from real-world use” [3]. That is a vendor telling you, in its own release documentation, that the ceiling moves and that routing down is a normal thing to do. If the company that built the model treats substitution as ordinary, a workflow of yours that cannot survive substitution is your design choice rather than theirs.
The practical version is a habit. When something you rely on stops behaving as it did, check the model version before you rewrite the prompt. Half the time you are debugging a policy decision, and no amount of prompt work will fix it.
Models leave on a schedule too
The arrival gate gets the attention. The departure gate is the one that will actually take a working automation off you, and unlike the arrival it comes with dates published months ahead.
Anthropic “notifies customers with active deployments for models with upcoming retirements, providing at least 60 days’ notice before model retirement for publicly released models” [5], and sorts models into Active, Legacy, Deprecated and Retired, where a retired model means “requests to retired models will fail” [5]. That policy produced a busy 2026. Claude Haiku 3 was announced on 19 February 2026 and retired on 20 April 2026; Claude Sonnet 4 and Claude Opus 4 were announced on 14 April 2026 and retired on 15 June 2026; Claude Opus 4.1 was announced on 5 June 2026 and retired on 5 August 2026 [5]. Each notice named a recommended replacement, which is generous and does not help at all if nobody read the notice.
OpenAI’s stated minimum is longer, and carries a clause worth noticing: “Unless safety or compliance concerns require a faster timeline, we provide the following minimum notice periods before model retirement: Generally available models: At least 6 months” [6]. Older GPT-5 and o3 snapshots were announced on 11 June 2026 for shutdown on 11 December 2026, and Sora 2 video models were announced on 24 March 2026 for shutdown on 24 September 2026 [6].
The exception clause is the whole point of this section. The same safety machinery that holds a capability back at the front of its life can shorten its notice period at the end. A model version is a rental with a lease term you did not negotiate, and the term is published on a page you have probably never opened.
Work that survives a model you cannot get
The fix is not vendor neutrality as an ideology, which usually means building an abstraction layer nobody maintains. It is much smaller than that.
Name the model in one place. If a version string appears in four automations, a saved prompt, a Cursor config and a Zapier step, a swap is an afternoon of archaeology. If it appears once, a swap is a line.
Keep a fallback you have actually run. A named alternative you have never sent real work through is a guess. Run one genuine week of output through the second choice, note what got worse, and write that down; the note is the thing that lets you decide in ten minutes rather than ten hours.
Keep five to ten real inputs with a short description of what a good answer looks like. This is the cheapest thing in the guide and the one most often skipped. Without it, evaluating a replacement model is vibes, and vibes take a week. With it, a swap is testable in an hour.
Then price the swap, because the number is usually smaller than the dread.
Places × minutes each. Computed in the page; nothing is sent anywhere.
If that number is uncomfortable, the problem is the first item on the list rather than the model. Nine places is an afternoon. Forty places is a project, and forty places is what accretion produces when nobody is counting.
What still goes wrong
Every date, quotation and policy statement above was read from its source on 5 September 2026, and this is a fast-moving area where the specifics move faster than the pattern. Executive orders are revoked and rewritten by later ones. Capability frameworks get new versions; Anthropic’s was on version 3.4 as of 8 July 2026 [4], which tells you how often these documents change. Treat the numbers here as a snapshot to re-check against the linked pages rather than as standing facts.
The larger honesty is that none of this preparation makes a capability appear. If the thing you need sits behind a trusted-access route you do not qualify for, no amount of tidy configuration changes that, and the correct answer is sometimes to build the smaller version of the product with what you can actually buy. Substitution has a floor. Routing down from a model rated High to one that is not is a genuine capability loss on some tasks, and pretending otherwise produces worse work rather than resilient work.
There is also a quieter failure that the checklist cannot catch. Most small teams do not have a model dependency they chose; they have one that formed. Somebody tried a model on a Tuesday, it worked, and the workflow set around it like concrete. The list above assumes you can find those places. In practice the ones that bite are the steps nobody remembers building, in a tool nobody thinks of as an AI tool, and you tend to find them on the morning the request starts failing.
- 01The White House — Executive Order 14409, Promoting Advanced Artificial Intelligence Innovation and Securitywhitehouse.gov
- 02OpenAI — GPT-6 Astra system card, Deployment Safety Hubdeploymentsafety.openai.com
- 03OpenAI — GPT-5.6 system card, Deployment Safety Hubdeploymentsafety.openai.com
- 04Anthropic — Responsible Scaling Policyanthropic.com
- 05Anthropic — Model deprecationsplatform.claude.com
- 06OpenAI — Deprecationsdevelopers.openai.com
- 07OpenAI — Trusted Access for Cyberopenai.com
- 08Fortune — Sam Altman says OpenAI made many changes during talks with US officialsfortune.com