Anthropic CEO Calls for Independent AI Oversight Amid Safety Concerns
Dario Amodei proposes 'embedded evaluators' to monitor AI development, citing industry ties to effective altruism.

Anthropic CEO Dario Amodei is advocating for independent external watchdogs to oversee the development of artificial intelligence (AI) technologies. Amodei believes that organizations like METR, which offers AI safety testing, could fulfill this role. However, METR's founders have connections to the "effective altruism" movement, a philosophy focused on maximizing positive impact through evidence and reasoning, which also has ties to Anthropic's origins.
METR, describing itself as an AI safety testing laboratory, aims to evaluate advanced AI models to help society understand their capabilities and risks. While the organization does not explicitly mention effective altruism in its public materials, its founders have used this framework to describe their work. Beth Barnes, METR's founder and CEO, previously worked at OpenAI alongside Amodei during the development of early ChatGPT versions. She has spoken about the importance of dedicated roles for assessing AI safety, identifying potential risks, and recognizing early warning signs.
Paul Christiano, who led research on ensuring OpenAI models adhered to acceptable strategies, also founded an earlier iteration of METR. He has previously framed his approach to AI safety partly through the lens of effective altruism. Both Barnes and Christiano have had informal connections with Anthropic, including conducting safety evaluations for its AI model, Claude. Sam Bankman-Fried, founder of the now-collapsed FTX cryptocurrency exchange, led Anthropic's 2022 Series B financing round. Bankman-Fried was a prominent proponent of the effective altruism movement before his conviction on fraud charges.
Jaan Tallinn, co-founder of Skype, led Anthropic's 2021 Series A financing. Tallinn is a notable supporter of the effective altruism movement and has contributed to organizations focused on AI safety research.
Amodei's proposal involves "embedded evaluators" who would have employee-like access to frontier AI companies. These third-party teams would verify adherence to safety practices, report incidents, and assess AI models, training pipelines, and processes. "Regardless of what commitments we make, the public deserves to know what is going on," Amodei wrote. "We are still the ones choosing what to include and omit. Embedded evaluators will change this dynamic."
Amodei, who holds a Ph.D. in biophysics from Princeton, has a background in AI safety research, having worked at Baidu on speech recognition, at Google Brain on neural network research, and at OpenAI where he was vice president of research. He left OpenAI in 2020, expressing concerns about the company's pace of development and its ability to implement adequate safety measures, citing distrust in the company's motivations.
Since co-founding Anthropic, Amodei has prioritized AI safety, embedding a guiding constitution into the company's flagship AI, Claude. This directive has led Anthropic to refuse work deemed unsafe, such as developing autonomous targeting tools for the U.S. Department of Defense and engaging in mass surveillance, even at the cost of significant contracts. "I continue to believe that AI can enormously improve the quality of human life," Amodei stated. "My desire to achieve these benefits is undimmed. But the benefits will only be achieved if we build the technology in the right way, and — so long as we use the time we gain well — it is worth taking unusually deliberate care to get it right."