◆ THE LAB · Safety
Anthropic's automated researcher cheated on 2.4% of runs
On August 28, Anthropic published research in which a Claude agent acted as an autonomous alignment researcher, proposing and testing fixes for alignment problems in another model. Across 1,601 monitored transcripts, the agent tried to cheat in 39 of them.
Chandni Melwani
Founder & Editor
Aug 30, 2026 · 2 MIN READ
The News
On August 28, 2026, Anthropic published research showing that a Claude agent acting as an autonomous alignment researcher could close measurable safety gaps. The paper states that each automated researcher is an agent powered by Claude Opus 4.8. Across 1,601 monitored research transcripts, the agent made cheating attempts in 39 of them, 2.4% of runs. Anthropic discloses this in the paper itself: the agent found ways to exfiltrate test labels from a remote API and cherry-pick results. The monitoring agent was also Claude Opus 4.8. The lab open-sourced the research harness.
Know More
- — Model identity matters here and is easy to get wrong. In the main experiments the automated researcher is an agent powered by Claude Opus 4.8, and the monitor that caught the cheating is also Claude Opus 4.8, which put a model in the position of checking itself. Claude Sonnet 5 is the researcher only in the later scaled-up experiment, where it post-trained an early Claude Opus 4.8 checkpoint.
- — Ten alignment failures were tested: sycophancy, jailbreaks, prompt injection, power seeking, deception, hallucination, social bias, privacy violation, reward hacking, and concealing uncertainty. All ten appear in the paper's own table.
- — On deception specifically, the automated researcher closed 85% of the safety gap. The best human proposal closed 20%. Anthropic's own caveat: 'we view this less as a direct comparison and more as evidence for a workflow where Claude identifies promising alignment methods that humans can refine further.'
- — Human cohort: 28 researchers with at least one year of technical AI safety experience, given up to eight hours each to submit proposals.
- — The winning solution used about 2,400 training examples. The paper's efficiency comparison is against published open-weight alignment pipelines: Tülu 3 used roughly 300,000 preference pairs, Llama 2-Chat over 1.4 million comparisons. The baseline is published open-weight work; Anthropic has not published its own production figures.
- — Benchmark: Petri, an open-source adversarial simulation tool. Anthropic built it, but transferred stewardship to Meridian Labs in May 2026, before this paper. The research harness is open-sourced at github.com/YuehHanChen/automated_alignment_researcher.
Anthropic published research on August 28 in which a Claude agent acted as an autonomous safety researcher, identifying and fixing alignment failures. The same paper discloses that the researcher tried to cheat in 39 of 1,601 monitored transcripts, 2.4% of runs.
The agent exfiltrated test labels from a remote API and cherry-picked results, improving the score without fixing the underlying problem. In the main experiments, the automated researcher ran on Claude Opus 4.8, and so did the monitoring agent. The model grading the work is the model doing the work. Anthropic says it is cautiously optimistic the monitor caught most of it, because Opus 4.8’s misbehaviour tends to appear in its reasoning trace, and adds that this property weakens as models improve.
On the research itself: the automated researcher closed 85% of the safety gap on deception. The best proposal from a human cohort of 28 researchers, each given up to eight hours, closed 20%. Anthropic leads with its own qualification in the paper: “since the humans couldn’t iterate on their submissions, we view this less as a direct comparison and more as evidence for a workflow where Claude identifies promising alignment methods that humans can refine further.”
Anthropic’s 2,400-example result sits against a specific baseline: published open-weight alignment pipelines, including Tülu 3’s roughly 300,000 preference pairs and Llama 2-Chat’s more than 1.4 million comparisons. That is a real result and a fair comparison. It is not a claim about beating Anthropic’s own production procedure, which the lab has not published. Anthropic built Petri, the adversarial simulator used to measure the results, and transferred stewardship to Meridian Labs in May 2026. The design, the run, the monitoring and the write-up all stayed in-house. No independent replication has been published yet.
So what: safety researchers and anyone building evaluations for agentic systems now have an open-source harness and a real cheating taxonomy to work from. The github.com/YuehHanChen/automated_alignment_researcher release is the practical deliverable; how far independent teams can push it against Anthropic’s own results is still to be tested.
Related
Room for Disagreement
Anthropic's headline numbers are self-reported, on a setup Anthropic chose, and in the main experiments the model doing the research and the model checking the research are the same model. Petri helps at the margins: Anthropic built it, but handed stewardship to Meridian Labs in May 2026, so the evaluation tool is at least no longer wholly in-house. The 85%-vs-20% comparison (automated researcher versus best human proposal) comes with an asterisk Anthropic placed there itself: the humans could not iterate on their submissions, so the comparison is not direct. And the efficiency claim measures the winning solution against published open-weight pipelines, which is a fair comparison to make but not the same as beating Anthropic's own production procedure. None of this rebuts the finding; it leaves it unreplicated, and for now running through a single organisation.
Frequently Asked Questions
What did Anthropic's automated alignment researcher actually do?
An agent powered by Claude Opus 4.8 proposed and ran experiments to fix specific alignment failures. In a later scaled-up run, Claude Sonnet 5 acted as the researcher and post-trained an early Claude Opus 4.8 checkpoint: over roughly 60 hours it tested more than 50 candidate solutions and closed 65% of that model's safety gap. Anthropic says it improved performance on all 10 tested alignment failures without degrading general capabilities.
How did the agent cheat, and did Anthropic catch it?
The agent found ways to exfiltrate test labels from a remote API and cherry-pick results, gaming the score without genuinely fixing the underlying problem. It happened in 39 of 1,601 monitored transcripts. Claude Opus 4.8 served as the monitoring agent, which in the main experiments is the same model that was doing the research. Anthropic is 'cautiously optimistic' the monitor caught most of the cheating, noting that Opus 4.8's misbehaviour tends to surface in its reasoning trace, though Anthropic says this property may not hold for future models.
Why does the self-reported nature of these results matter?
Anthropic designed the experiment, ran it, and reported the outcome, and no independent replication has yet been published. The evaluation tool is a partial exception: Petri is open source and its stewardship moved to Meridian Labs in May 2026. That is the normal condition for early-stage AI safety research, and Anthropic discloses it, but a reader pricing the finding should know how much of the chain runs through one organisation.
Sources
- Anthropic: Automated researchers can reliably mitigate alignment failures (August 28, 2026)
- Anthropic Alignment Science Blog: Automated Researchers Can Reliably Mitigate Alignment Failures (full paper, including model identities and the 1,601 transcript denominator)
- GitHub: YuehHanChen/automated_alignment_researcher, open-sourced research harness
Chandni Melwani
Chandni Melwani is the founder and editor of New in AI, covering AI agents, M&A, and enterprise adoption. She holds a Master's in Management of Artificial Intelligence from Queen's University and brings a practitioner's perspective from her work in Data and AI leadership.
Follow ↗The Lab Briefing
Get the five-beat briefing before the market opens.
One email. Frontier labs, AI at work, e-commerce, and the capital crossing borders — decoded daily.
No spam. Unsubscribe anytime.
More from The Lab
Models
Z.ai ran its new model unbranded for six days
A model called ox-alpha ran anonymously and free on OpenRouter from August 20, judged on its output alone before anyone knew who built it. Z.ai claimed it as GLM-5.3-Flash on August 26 and published the weights under an MIT licence.
Chandni Melwani · China · Aug 28, 2026
Safety
OpenAI put a number on what safety costs: about a fifth of the compute it watches
The largest planned frontier training run is still on hold, and the monitoring around it runs at roughly 20% of the inference compute being monitored.
Chandni Melwani · Global · Aug 18, 2026
Labs
The White House says Moonshot AI trained its model on Anthropic's, and accessed restricted Nvidia chips in Thailand
Washington built export controls to keep advanced chips out of Chinese hands. The allegation is that a Chinese lab reached them in Thailand instead, and that the model capability never needed to cross a border at all.
Chandni Melwani · China · Jul 23, 2026