Technology
Anthropic and OpenAI back embedding third-party safety evaluators in labs
Leaders propose placing external researchers inside AI companies, though critics note the monitors lack enforcement authority.
The short version
- Anthropic CEO Dario Amodei and OpenAI CEO Sam Altman proposed embedding outside safety evaluators inside frontier AI labs to inspect alignment and log safety incidents.[TechCrunch · CNBC]
- Experts question the model's independence, noting embedded observers lack regulatory enforcement power or the ability to halt model releases.[CNBC]
- Previous outside safety reviews, including Apollo Research's review of Astra, faced severe time constraints that hindered confident conclusions.[TechCrunch]
Key facts
- Anthropic CEO Dario Amodei proposed placing third-party safety monitors inside frontier AI firms to evaluate model alignment and disclose safety incidents.[TechCrunch · CNBC]
- OpenAI CEO Sam Altman stated that OpenAI would also commit to hosting embedded third-party evaluators.[TechCrunch]
- Banking regulation scholar Julie Andersen Hill observed that unlike statutory bank examiners, proposed AI evaluators lack the legal authority or kill switch to stop development or releases.[CNBC]
- Safety organization METR stated that it accepts neither cash payments nor donations from artificial intelligence companies or their executives.[CNBC]
- Apollo Research received only three days to evaluate the Astra model, limiting its ability to establish firm findings.[TechCrunch]
What remains uncertain
- It remains disputed whether voluntary embedded evaluation programs can deliver objective oversight without statutory backing or independent enforcement power.[TechCrunch · CNBC]
Sources
Outlet counts describe coverage, not independent confirmation. Reports may share a wire service or original source.