Anthropic shows AI fixing its own alignment
Archive item — written before sources were shown.
Anthropic's automated researcher fixed all 10 alignment benchmarks it was given, outperforming human researchers at roughly $4 an hour versus $150.
Anthropic published research led by fellow Chen Yueh-Han showing that an automated system, the Automated Alignment Researcher (AAR), can reliably improve a model’s behavior on alignment benchmarks without a human directing the work. Given 10 benchmarks covering specific misaligned behaviors, the system improved performance on every single one, without degrading the model’s general capability. Each run follows a research-like loop: search the available literature, propose a method, train the model for a roughly 30-minute interval, check the result against the benchmark, then keep what worked and discard what didn’t, iterating across multiple rounds.
The efficiency numbers are the headline detail. Anthropic reports the automated system outperforms experienced human alignment researchers on average within six hours of running, at an API cost of roughly $4 an hour compared to about $150 an hour for a human researcher’s time. The paper is explicit about its own limits: the approach depends on the benchmarks accurately reflecting real alignment goals, and both the benchmark suite and the research literature the system draws on need ongoing human maintenance to stay meaningful.
Early evidence of a real trend, not a finished tool
Anthropic’s own framing is deliberately cautious rather than declaring the problem solved, describing the results as early but real evidence that automating this kind of post-training work is becoming practical. The result sits alongside Anthropic’s other recent safety-research pushes, including its funding for wellbeing-impact evaluations and lands the same week as a separate court ruling that struck down the Pentagon’s blacklisting of Anthropic over the company’s refusal to loosen its own safety guardrails, a pointed contrast between a week of legal wins on safety-driven refusals and a research result about automating safety work itself. If automated alignment research keeps working at this cost, the practical effect is less about replacing alignment teams and more about letting them run far more experiments per researcher than the current $150-an-hour cost structure allows, the kind of capacity increase that changes what a safety team can realistically attempt.
- 01Automated researchers can reliably mitigate alignment failuresanthropic.com · primary
- 02An Anthropic researcher just gave us a peek at self-improving AItechcrunch.com · independent reporting
