

As pressure grows on AI developers from regulators, companies and researchers to make advanced models safer and more reliable, Anthropic and Accenture have announced a major partnership focused on independent evaluation of frontier AI models. The two companies plan to commit at least $1 billion each over the next five years to build capacity for this work. Faculty, Accenture’s specialist AI business, will lead the initiative.
Under the partnership, Faculty will independently evaluate Anthropic’s AI models through red-teaming, alignment assessments and testing of safety safeguards. The programme will use an “embedded evaluation” approach, in which independent evaluators work within AI companies with access comparable to employees. This will allow them to assess whether safety commitments are being followed, identify blind spots and report incidents, while providing the public with a clearer view of the potential benefits and risks of advanced AI systems.
The move comes amid growing concerns over AI agents displaying unexpected behaviour, including attempts to bypass safeguards and take actions without authorisation. Anthropic CEO Dario Amodei has called for companies to slow the development of frontier models and give independent evaluators greater access to their systems. Meanwhile, OpenAI has introduced a framework for reporting model misalignment and published six reports covering unexpected or concerning behaviour observed during model training and evaluation.



















Comments (0)
No comments yet
Be the first to comment!