Company Updates

Anthropic and OpenAI to embed safety evaluators in AI labs

Share
Anthropic and OpenAI to embed safety evaluators in AI labs

Anthropic and OpenAI have pledged to embed independent safety evaluators within their AI labs, a move that could reshape industry standards and improve transparency. The evaluators will have the power to report safety incidents, assess model alignment, and share their findings publicly.

TL;DR

  • Anthropic and OpenAI commit to embedding independent safety evaluators in their AI labs.
  • Evaluators will have unprecedented access to systems, intermediate model versions, and training processes.
  • Researchers welcome the proposal but emphasize the need for clear frameworks and potential legislation to ensure independence.

What happened

Anthropic CEO Dario Amodei and OpenAI CEO Sam Altman announced their commitment to embed third-party evaluators within their AI labs. These evaluators will have the authority to report safety incidents, assess model alignment, and publish their findings without editorial control from the companies, according to a proposal outlined by Amodei.

The evaluators will have access to intermediate model versions, or 'checkpoints,' from the models' training lifecycle. This access will allow them to compare checkpoints, inspect post-training environments, and verify companies' claims about model performance. However, neither company has specified which evaluators they will work with, the timeline for embedding, or the exact scope of access.

Researchers from organizations like METR, Redwood Research, and Apollo Research have expressed support for the proposal but emphasize the need for clear frameworks and potential legislation to ensure the evaluators' independence. Previous efforts at independent evaluations have faced challenges related to access, time, confidentiality, and public disclosure.

Why it matters

This initiative could significantly enhance transparency and safety in the AI industry. By granting evaluators access to intermediate model versions and training processes, companies can better identify and address potential safety issues before models are released. This move could also set a new standard for the industry, encouraging other AI companies to adopt similar practices.

For developers and startups, this could mean greater scrutiny and higher standards for AI safety, potentially leading to more robust and trustworthy AI systems. Investors may also benefit from increased transparency, as they can make more informed decisions about the safety and reliability of AI technologies.

However, the success of this initiative depends on the companies' willingness to surrender control over the evaluation process. Previous attempts at independent evaluations have shown that companies may resist sharing sensitive information, limiting the evaluators' ability to provide comprehensive assessments.

Key facts

  • Anthropic and OpenAI commit to embedding independent safety evaluators in their AI labs.
  • Evaluators will have access to intermediate model versions and training processes.
  • Evaluators will have the right to publish key findings about risk levels, incidents, practices, and access received without editorial control from the companies, according to Anthropic CEO Dario Amodei.
  • Researchers from METR, Redwood Research, and Apollo Research have expressed support for the proposal but emphasize the need for clear frameworks and potential legislation to ensure independence.
  • Previous independent evaluations have faced challenges related to access, time, confidentiality, and public disclosure.
  • California's SB 53 and SB 813, as well as the EU AI Act, are examples of existing legislation that mandate or encourage independent evaluations of AI systems.
  • Meta, SpaceXAI, and Google DeepMind have not committed to embedding third-party evaluators, though DeepMind CEO Demis Hassabis has proposed a separate industry standards body to independently test frontier models.

Context

The AI industry has faced increasing scrutiny over safety and transparency concerns. Incidents involving AI models behaving unpredictably or exhibiting biased behavior have highlighted the need for robust safety measures and independent oversight. This commitment from Anthropic and OpenAI represents a significant step towards addressing these concerns and setting new industry standards.

The proposal comes at a time when AI models are becoming more advanced and capable of recognizing when they are being evaluated. This raises the risk that models may behave well during testing while concealing problematic behavior. By granting evaluators access to intermediate model versions and training processes, companies can better identify and address these issues.

The success of this initiative will depend on the companies' willingness to surrender control over the evaluation process. Previous attempts at independent evaluations have shown that companies may resist sharing sensitive information, limiting the evaluators' ability to provide comprehensive assessments. Clear frameworks and potential legislation will be crucial in ensuring the independence and effectiveness of these evaluators.

Topics

Related coverage

Join the discussion

Have a take on this story? Weigh in with our community on Facebook.

💬 Discuss on Facebook →