tuesday, october 6, 2026 · the day's ai, attributed published by trilot llc · wyoming
guide · judgment & safety

The capability tier your model ships with

Read a lab's capability classification the way you read a service tier, and set your work up so a refusal or a gate slows you down instead of stopping you.

Published 2026-09-05 · Updated 2026-09-05 · Read 10 min · Reviewed by Rami Steitieh

Verified 2026-09-05 · Rami
on this page · 0 / 0 checked

A model you have been paying for stops doing something it did last month. Same prompt, same account, and the refusal is polite and useless. Or a new flagship launches, you go to switch it on for the team, and it is not in the list, because somebody has to enable it first. Neither of those is a bug, and neither is a jailbreak story. Both are the visible end of a decision the lab made before the model shipped, about how capable it judged the model to be and what it was prepared to let that capability do in your hands.

Frontier labs now grade their own models against written capability thresholds and attach safeguards to the grade. The grading is not marketing. It decides what the product will do, who can turn it on, and how quickly the restrictions come off. This guide is for someone running a small business or a solo practice on top of these models who wants to read that grading the way they read a service tier. It is not for safety researchers, who need the evaluation literature rather than a summary. It is also not a guide to getting around a refusal, which is a poor use of an afternoon and leaves you depending on a gap the vendor intends to close.

A threshold is a decision rule, not a warning label

OpenAI’s Preparedness Framework, updated on 15 April 2025, tracks three categories it calls “established areas where we have mature evaluations and ongoing safeguards”: biological and chemical, cybersecurity, and AI self-improvement [2]. Each has two thresholds. High capability covers a model that could amplify existing pathways to severe harm, and requires safeguards before deployment. Critical capability covers a model that could introduce unprecedented new pathways to severe harm, and requires those safeguards during development as well [2].

The cyber threshold is written tightly enough to argue with. A model reaches Critical, OpenAI says, “if it can identify and develop functional zero-day exploits of all severity levels in many hardened real-world critical systems without human intervention, or can devise and execute end-to-end novel strategies for cyberattacks against hardened targets given only a high level desired goal” [1]. That is a much higher bar than writing exploit code when a person walks it through the steps.

Here is the part that matters for anyone planning around these systems. The trigger is not proof. On 7 August 2026, writing about a then-unreleased model called Astra, OpenAI said its preliminary evaluations showed performance strong enough that “we cannot rule out Critical capability level at this time” [1]. It then paused “internal activities involving Astra that do not yet meet these strengthened security control requirements” [1]. Nobody had demonstrated the model was Critical. The lab could not demonstrate it was not, and that was sufficient to act.

Anthropic reached the same construction more than a year earlier. When it activated ASL-3 protections for Claude Opus 4 on 22 May 2025, it said it had “not yet determined whether Claude Opus 4 has definitively passed the Capabilities Threshold that requires ASL-3 protections”, and that it could not clearly rule out those risks, so it applied the protections anyway [5]. Two labs, two rubrics, one grammar: absence of a clean negative is treated as a positive.

What a Critical classification did to a shipping product

Astra shipped. Its system card, published on 3 September 2026, calls it “our first model to reach the Critical level of cybersecurity capability under our Preparedness Framework” [3]. The launch post says the same thing more plainly: “Astra is a significant jump in cyber capabilities and meets the Critical threshold in cybersecurity under our Preparedness Framework” [4].

So the classification did not mean the model was withheld. It meant the model arrived with a capability surface smaller than the model itself. OpenAI is specific about the shape of that surface: “defenders can use it to complete tasks such as secure code review and patching. However, Astra will refuse to comply with more advanced cybersecurity tasks such as creating proof-of-concept exploits for vulnerabilities” [4].

The second gate is commercial rather than behavioural, and it arrives in stages. Astra is “rolling out today to a limited set of organizations and over the coming days will become available to all ChatGPT Plus, Pro, Business, and Enterprise users, as well as through the OpenAI API, Microsoft Azure, and AWS Bedrock” [4]. It lists at $10 per million input tokens and $50 per million output tokens, with a fast mode in the API that “delivers up to 2x the speed of Standard processing at 2x the Standard price” [4]. On Enterprise, “access is off by default at launch”, and an administrator has to enable it for the workspace [4]. The restriction also has an expiry that is not a date. Through a programme OpenAI calls Daybreak, it says it plans “to expand access and roll out less restrictive safeguards in the coming weeks” [4].

The third gate is aimed at accounts rather than tasks. For users flagged as potentially high risk, the system card says OpenAI has “additionally trained in the ability to adjust the model’s refusal boundary to be more conservative”, covering a broader range of dual-use risks [3]. Two accounts on the same plan, sending the same words, can get different answers.

Three labs, one machine

It would be convenient if this were one vendor’s quirk, because then the fix would be a vendor change. It is not.

Anthropic’s Responsible Scaling Policy, at version 3.4 effective 8 July 2026, defines AI Safety Levels with capability thresholds attached, including thresholds for AI research and development and for chemical, biological, radiological and nuclear development [6]. Its ASL-3 Deployment Standard is a layered defence built from access controls tailored to the deployment context, real-time prompt and completion classifiers, asynchronous monitoring and post-hoc jailbreak detection, while the matching Security Standard spans 17 areas including access management, software supply chain security, and red teaming and penetration testing [6]. When Anthropic actually switched that on in May 2025, it meant real-time classifiers over model inputs and outputs, a bug bounty pointed at those classifiers, more than 100 security controls, two-party authorisation for model weight access, and egress bandwidth controls to make exfiltration harder [5].

Google DeepMind publishes a Frontier Safety Framework, in its third iteration dated 22 September 2025 and updated on 17 April 2026, organised around Critical Capability Levels, defined as “capability levels at which, absent mitigation measures, frontier AI models or systems may pose heightened risk of severe harm” [7]. The commitment attached is procedural: “we conduct safety case reviews prior to external launches when relevant CCLs are reached” [7].

The vocabulary differs and the structure does not. Write the thresholds down in advance, evaluate against them, and let the result decide what ships and to whom. That convergence is the durable fact in this whole topic. Moving from one lab to another does not move you out of the mechanism, and the lab you move to may have its own model in a gated state on a different week.

Telling a refusal apart from a limit

When a model will not do something, there are three ordinary explanations, and the fixes have nothing in common, so spending 5 minutes on the diagnosis pays for itself.

The first is that the model cannot do it. Run the same request against a cheaper model in the same family. If both fail the same way, producing something confused or wrong rather than declining, you are looking at a capability limit, and better context or a stronger model is the lever.

The second is that the classification gates it. A policy refusal usually reads as unwilling rather than incapable, arrives immediately, and is stable no matter how you rephrase. Check the model’s system card and the vendor’s usage policy for whether your task sits in a gated category. If it does, the answer is a different supplier or a properly authorised route, not a cleverer prompt. The Astra case shows how narrow the line can be: secure code review and patching are in, proof-of-concept exploits are out, and both of those are things a security-minded developer does in the same afternoon [4].

The third is that your plan or your administrator has it switched off. This is the least interesting and the most common, and it is indistinguishable from the model not existing. Astra on Enterprise is the worked example, off until somebody with admin rights turns it on [4].

Build so a tier change is a config change

The realistic exposure for a small operator is not that a lab withholds a model. It is that a capability you built on narrows, widens, or moves to a different plan on a schedule set by someone else, and your work either bends or stops.

Three habits cover most of it. Name the model version in exactly one place, a config value or an environment variable, rather than scattering it through prompts, scripts and automations where nobody can find all of them under pressure. Keep a second vendor’s model configured and actually tested against your real work, not merely signed up for, because an untested fallback is a plan you have never run. And know the price of the tier below the one you are on before the day you need it.

That last one is arithmetic. OpenAI’s pricing page lists GPT-6 Astra at $10 per million input tokens and $50 per million output, and the next model down the list, GPT-5.6 Sol, at $4 and $20, with a note that Sol’s promotional pricing is available at least through 21 November 2026 [8]. The premium you pay for the top tier is worth holding as a monthly number rather than a feeling.

calculator
What the top tier costs over the one below
— $ / month

Astra list rates minus GPT-5.6 Sol, at the prices cited above. Computed in the page; nothing is sent anywhere.

If that number is small, staying on the flagship is easy to justify and a forced downgrade is survivable. If it is large, the useful question is which of your tasks genuinely needs the top tier, and the honest answer is usually fewer of them than you assumed. Either way you have made the decision once, in calm conditions, instead of at 9am on the morning a gate moves.

checklist
When a model refuses or a tier moves
0 of 8 · saved in this browser only

What still goes wrong

The controls get announced at a level of detail you cannot plan against. Astra’s system card lists Trusted Access for Cyber and Actor Level Enforcement among its safeguards without setting out who qualifies, how to apply, or what a flagged account loses [3]. The relaxation schedule is “in the coming weeks” [4]. If part of your business depends on a specific capability being unlocked by a specific date, no published document currently supports that plan, and no amount of reading will change that.

The grading is also done by the party with the strongest commercial interest in the answer, against a rubric that party wrote and keeps revising. OpenAI’s framework was updated in April 2025 [2], Anthropic’s policy is on version 3.4 as of July 2026 [6], and Google’s is on a third iteration updated in April 2026 [7]. Revision is not evidence of bad faith, since the underlying science is young and evaluations genuinely improve. It does mean a threshold is a commitment about process rather than a guarantee about outcomes, and that the same model might land in a different band under a different lab’s rubric.

Most importantly, none of this addresses the failure that will actually cost you money. Capability thresholds concern severe, low-frequency harms in cyber, biology and self-improvement. They say nothing about a model inventing a figure in a client report on a Thursday. A model graded below Critical in every category can still be wrong in a way that loses you a customer, and a model graded Critical is no more reliable about your invoice than the one it replaced. If you want a signal about everyday accuracy, the classification is the wrong instrument entirely. And if you run a security practice that needs offensive tooling, this guide will not get you access; the vendors’ trusted access programmes decide that, on their terms and their timetable.

sources
  1. 01OpenAI — Responding to the next frontier of critical cyber capabilitiesopenai.com
  2. 02OpenAI — Updating our Preparedness Frameworkopenai.com
  3. 03OpenAI — GPT-6 Astra system card (Deployment Safety Hub)deploymentsafety.openai.com
  4. 04OpenAI — Introducing GPT-6 Astraopenai.com
  5. 05Anthropic — Activating ASL-3 protectionsanthropic.com
  6. 06Anthropic — Responsible Scaling Policyanthropic.com
  7. 07Google DeepMind — Strengthening our Frontier Safety Frameworkdeepmind.google
  8. 08OpenAI — API pricingdevelopers.openai.com
next guide
Hidden reasoning is not a security boundary
9 min · verified 2026-09-05
related guides