

Anthropic has reported early evidence that AI systems can increasingly take on parts of the research process used to improve other AI models. In a study published on August 28, the company developed automated alignment researchers (AARs) powered by Claude Opus 4.8. These systems searched research literature, proposed training methods and iteratively trained target models to address 10 different alignment failures, including deception, sycophancy and jailbreak-related behaviour. Anthropic said the methods improved the targeted safety benchmarks without degrading measured general capabilities.
The study also compared the automated systems with proposals from 28 experienced human AI-safety researchers. Anthropic reported that its best automated methods outperformed the human proposals after iterative testing. In a separate experiment, Claude Sonnet 5 was used to improve an early Claude Opus 4.8 checkpoint, reaching alignment scores close to the released model with a much smaller training set. Researchers, however, stressed that the findings have important limitations: the experiments covered only selected, measurable alignment failures, and benchmark improvements do not necessarily prove real-world alignment. Anthropic described the results as early positive signals, rather than evidence that AGI has already been achieved.



















Comments (0)
No comments yet
Be the first to comment!