

Anthropic has published research examining whether AI models can take a larger role in researching and improving the safety of other AI systems. The company tested Claude on 10 types of alignment failures, including deception, sycophancy, reward hacking and privacy violations. Claude was able to study existing research, propose training methods and datasets, train target models and evaluate the results. It could also use the findings from one experiment to develop and test another approach.
In one experiment, Claude significantly reduced the safety gap in deception tests and outperformed human researchers under the specific experimental conditions. Anthropic also tested whether a less capable Claude model could develop alignment training for a more capable model, with the results showing that it could produce useful methods. However, the researchers noted several limitations, including the small number of alignment problems studied and the difficulty of monitoring increasingly capable AI systems. Anthropic stressed that the research does not demonstrate full self-improving AI, but it provides an early indication that AI could become more involved in future AI safety research.



















Comments (0)
No comments yet
Be the first to comment!