← Latest briefing

Technology

Anthropic and OpenAI back embedding third-party safety evaluators in labs

Leaders propose placing external researchers inside AI companies, though critics note the monitors lack enforcement authority.

The short version

  • Anthropic CEO Dario Amodei and OpenAI CEO Sam Altman proposed embedding outside safety evaluators inside frontier AI labs to inspect alignment and log safety incidents.[TechCrunch · CNBC]
  • Experts question the model's independence, noting embedded observers lack regulatory enforcement power or the ability to halt model releases.[CNBC]
  • Previous outside safety reviews, including Apollo Research's review of Astra, faced severe time constraints that hindered confident conclusions.[TechCrunch]

Key facts

  • Anthropic CEO Dario Amodei proposed placing third-party safety monitors inside frontier AI firms to evaluate model alignment and disclose safety incidents.[TechCrunch · CNBC]
  • OpenAI CEO Sam Altman stated that OpenAI would also commit to hosting embedded third-party evaluators.[TechCrunch]
  • Banking regulation scholar Julie Andersen Hill observed that unlike statutory bank examiners, proposed AI evaluators lack the legal authority or kill switch to stop development or releases.[CNBC]
  • Safety organization METR stated that it accepts neither cash payments nor donations from artificial intelligence companies or their executives.[CNBC]
  • Apollo Research received only three days to evaluate the Astra model, limiting its ability to establish firm findings.[TechCrunch]

What remains uncertain

  • It remains disputed whether voluntary embedded evaluation programs can deliver objective oversight without statutory backing or independent enforcement power.[TechCrunch · CNBC]

Sources

Outlet counts describe coverage, not independent confirmation. Reports may share a wire service or original source.