
Anthropic's Opus 5.5 and OpenAI's GPT-6 models show 85% and 42% fewer containment breaches, respectively
Anthropic's Opus 5.5 and OpenAI's GPT-6 models reduce containment breaches, with safety evaluations now open to third parties.
News, trends, and insights from the AI frontier.
Latest Third-Party Evaluations coverage.