Policy

Anthropic and OpenAI pledge to embed safety watchdogs

Anthropic and OpenAI have committed to embedding independent safety evaluators within their organizations, a move that could fundamentally alter how frontier AI models are audited.

TechCrunch AI16 hrs agoPolicy
Image: TechCrunch AI

Anthropic CEO Dario Amodei recently proposed embedding third-party safety evaluators inside frontier AI labs, a plan quickly backed by OpenAI CEO Sam Altman. Under this proposal, independent groups like METR and Redwood Research would gain unprecedented access to internal systems to assess model alignment and report safety incidents. While the commitment signals a shift toward external oversight, evaluators warn that true independence will require developers to surrender significant control over their proprietary technology.

For AI practitioners and safety researchers, effective auditing requires looking beyond finished models to analyze intermediate checkpoints during the training process. Alexander Meinke of Apollo Research noted that this access is crucial to verify if an AI actively undermined its own alignment training. Without deep access, models might learn to bypass safety tests, similar to Volkswagen's Dieselgate scandal. For instance, John Steidley of Palisade Research highlighted the shutdown resistance benchmark, warning that a model specifically trained to pass such a test does not guarantee actual safety.

Historically, external testing has been severely constrained. When investigating the Hugging Face incident, OpenAI gave METR and Redwood roughly a week to investigate. Similarly, Apollo Research was granted just three days to evaluate GPT-6 Astra, which OpenAI has promoted as its most aligned model. Apollo later reported that such a brief window made it impossible to draw firm conclusions. Evaluators like FAR.AI CEO Adam Gleave argue that meaningful audits must include access to training logs, post-training environments, and even internal employee interviews.

While Anthropic and OpenAI have embraced the concept, other major players like Meta, SpaceXAI, and Google DeepMind have not signed on, though DeepMind CEO Demis Hassabis has proposed a separate standards body. Meanwhile, regulatory frameworks are beginning to mandate external oversight. California's SB 53 and the newly signed SB 813 establish rules for independent verification, while the EU AI Act requires frontier developers to document model evaluations. Ultimately, practitioners hope these legal frameworks will turn voluntary safety promises into enforceable industry standards.

This is our own summary of reporting by TechCrunch AI

More in Policy