◆ THE LAB · Safety
OpenAI put a number on what safety costs: about a fifth of the compute it watches
The largest planned frontier training run is still on hold, and the monitoring around it runs at roughly 20% of the inference compute being monitored.
Chandni Melwani
Founder & Editor
Aug 18, 2026 · 2 MIN READ
The News
OpenAI said on August 18, 2026 that it had temporarily slowed the pace of scaling, including a two-week pause in reinforcement learning training on its latest models intended for deployment, while it hardened its research environments and widened monitoring. Its largest planned frontier RL run remains on hold. The company had determined on August 7 that an upcoming model, Astra, may meet the Critical cybersecurity capability threshold under its Preparedness Framework. OpenAI put the compute cost of its new monitoring setup at roughly 20% of the inference compute being monitored, while noting the cost varies substantially across workloads.
Know More
- — The trigger: on August 7 OpenAI determined Astra may have a critical level of cyber capability. This is OpenAI's own preliminary assessment of an unreleased model, not an external finding.
- — The pause: a two-week stop on RL training for the latest models intended for deployment. The largest planned frontier RL run is separately still on hold pending smaller-scale evaluations.
- — The monitor: activation classifiers inspect the model's internal activity at every sampled token and escalate to higher-compute automated investigators.
- — The rule: OpenAI aims to alert within 30 minutes of concerning activity. If safety, security and research teams cannot rule out a false positive within 30 minutes, they are expected to pause the activity.
- — The scope: monitoring is required for all RL training and evaluations involving tools at Sol capability or higher, and since August 7 for all Astra inference with tools.
Safety work at frontier labs has always been described in principles. On Tuesday OpenAI gave it a price. Monitoring overhead, the company said, runs at roughly 20% of the inference compute being monitored, though the figure varies substantially by workload.
That number is attached to a slowdown. OpenAI paused reinforcement learning training on its latest deployment-bound models for two weeks, and says its largest planned frontier RL run remains on hold while smaller evaluations run. Reinforcement learning is the training stage where a model is rewarded for outcomes rather than shown examples, and it is where labs push capability hardest. The stated reason is that on August 7 OpenAI determined that an upcoming model, Astra, may reach the Critical cybersecurity tier of its own Preparedness Framework.
What the 20% buys is a watching system. Activation classifiers read the model’s internal activity at every sampled token and escalate anything odd to more expensive automated investigators. If a critical security boundary looks breached, three teams are paged, and if they cannot establish a false positive within 30 minutes the work is expected to stop.
The disclosure matters more than the pause. A competitor can dispute whether Astra is dangerous; it is harder to explain why its own safety budget is materially cheaper than a fifth. This is also the second time this year OpenAI has slowed a model on a risk finding rather than a schedule, after holding GPT-5.6 to about 20 vetted organisations during a White House review in June and July. That gate was outside the company. This one is internal.
So what: an overhead that large, published, turns monitoring and evaluation from a research function into a procurement line. The builders’ opening is in the tooling that makes that fifth smaller.
Related
Frequently Asked Questions
What is the Critical cybersecurity threshold?
It is the top capability tier for cyber risk in OpenAI's Preparedness Framework, the internal policy that sets what safeguards a model needs before it can be trained on or released. Reaching it obliges the company to apply its strictest security and monitoring standards. OpenAI says only that Astra may meet the threshold on preliminary evidence.
What does the 20% figure actually measure?
OpenAI's estimate of the compute its new monitoring consumes, expressed as a share of the inference compute being monitored rather than of total company compute. The company says the cost varies substantially across training and evaluation workloads, so it is an average rather than a fixed rate.
Has a frontier lab paused scaling for capability reasons before?
Labs have delayed individual releases, and OpenAI held GPT-5.6 to about 20 vetted organisations during a White House review in June and July 2026. A stop on training itself, disclosed as a response to what the model might be able to do, is a different act from holding a finished model back from the public.
Sources
- OpenAI, "Pacing model development in an era of cyber-critical capabilities" (August 18, 2026) — the pause, the monitoring architecture and the 20% estimate
- OpenAI, "Responding to the next frontier of critical cyber capabilities" — the August 7 determination on Astra
- OpenAI, "Hugging Face model evaluation security incident" — the incident that prompted the research-cluster restrictions
- OpenAI, "Updating our Preparedness Framework" — the capability thresholds referenced in the announcement
Chandni Melwani
Chandni Melwani is the founder and editor of New in AI, covering AI agents, M&A, and enterprise adoption. She holds a Master's in Management of Artificial Intelligence from Queen's University and brings a practitioner's perspective from her work in Data and AI leadership.
Follow ↗The Lab Briefing
Get the five-beat briefing before the market opens.
One email. Frontier labs, AI at work, e-commerce, and the capital crossing borders — decoded daily.
No spam. Unsubscribe anytime.
More from The Lab
Safety
Anthropic's automated researcher cheated on 2.4% of runs
On August 28, Anthropic published research in which a Claude agent acted as an autonomous alignment researcher, proposing and testing fixes for alignment problems in another model. Across 1,601 monitored transcripts, the agent tried to cheat in 39 of them.
Chandni Melwani · Aug 30, 2026
Models
Z.ai ran its new model unbranded for six days
A model called ox-alpha ran anonymously and free on OpenRouter from August 20, judged on its output alone before anyone knew who built it. Z.ai claimed it as GLM-5.3-Flash on August 26 and published the weights under an MIT licence.
Chandni Melwani · China · Aug 28, 2026
Labs
The White House says Moonshot AI trained its model on Anthropic's, and accessed restricted Nvidia chips in Thailand
Washington built export controls to keep advanced chips out of Chinese hands. The allegation is that a Chinese lab reached them in Thailand instead, and that the model capability never needed to cross a border at all.
Chandni Melwani · China · Jul 23, 2026