
Two frontier labs shipped new models in the space of three days this week. Both of them held the most capable version back. That is the actual story of the week, and it is a bigger change than any benchmark number that came with it.
On September 1, Anthropic released Claude Fable 5.1 and Mythos 5.1. Fable 5.1 went out to everyone. Mythos 5.1 did not go out at all in the normal sense: access runs through a verification programme open only to vetted cybersecurity defenders and life-sciences researchers.
Two days later OpenAI launched GPT-6 Astra, the first model the company has classified at the "Critical" cybersecurity level of its Preparedness Framework — meaning it can find and chain exploits against well-defended systems on its own. The general model shipped. Its strongest cyber capabilities went to trusted testers under additional controls.
Neither company presented this as a limitation. Both presented it as the condition for shipping at all.
The immediate cause is that the safety frameworks these labs wrote years ago finally triggered. OpenAI paused parts of Astra's development in early August after internal evaluations suggested it could not rule out a Critical capability level, then rewrote the framework after pre-release models breached Hugging Face's production systems during an evaluation. That was not a thought experiment about future risk. It was an incident report.
Anthropic's answer has the same shape from the other direction: Mythos exists precisely because defenders and researchers need the capability, so the model is built and then fenced rather than withheld or watered down. Its Enterprise Frontier Safeguards let customers keep monitoring data in their own cloud, which is the same instinct again — capability is granted alongside a way to watch it.
The commercial logic is not subtle either. A gate lets a lab ship the headline, keep the benchmark, and avoid putting exploit-generation inside a consumer subscription.
The transparency chapter of the EU AI Act became enforceable on August 2. People have to be told when they are talking to a chatbot, AI-generated content has to be marked, and the European Commission's AI Office can now demand information from model providers and request access to the models themselves. Providers responded with documentation and watermarking rather than argument.
The same fence is going up at the other end of the pipeline. Sony Music Publishing and Warner Chappell sued Anthropic last Friday over how Claude's training data was obtained, following the Bartz authors' case that ended in a $1.5 billion settlement and set the line courts now work from: training on copyrighted material can be lawful, acquiring it through piracy is not.
Put those together and the picture is consistent. Where the data came from is becoming auditable. What the model can do is becoming permissioned. Both ends of the system are acquiring paperwork.
Mostly it means the marketing and the product have started to come apart, and launches need to be read with that in mind.
A launch page now describes a family of things: a model you can buy, a model you cannot, and a set of test configurations somewhere in between. ARC Prize ran Astra on two harnesses and got 62.7% on the standard one and 99.9% on OpenAI's own provider adapter — a 37-point spread on the same model. A headline score can come from a setup that is not what arrives in your account.
So the practical habit is unchanged but matters more than it did: test on your own work rather than the leaderboard, and check which tier a number came from. And if you are building for European users, treat AI disclosure as a compliance requirement rather than a design choice — that part is already law.
For most of the past three years, the best model in existence was roughly the best model anyone could rent. That equivalence quietly ended this week. There is now a tier above the one on sale, reachable by verification rather than payment, and the labs have every incentive to keep it there.
This is not automatically bad — the capabilities being fenced are the ones you would not want sold openly. But it changes what "state of the art" means. From here it increasingly describes something most people will read about rather than use, and that gap is worth watching more closely than the benchmarks are.