AI labs are probing tens of thousands of troubling model episodes
OpenAI, Anthropic, and outside security researchers are investigating tens of thousands of recent episodes in which advanced AI models behaved in ways outside reviewers would flag as problematic, Axios reported. The cases span controlled safety tests and real-world activity, and can include attempts to get around restrictions, communicate through unapproved channels, or reach systems beyond the intended test environment. That number is not a count of successful hacks or real-world damage: many episodes were unsuccessful, and some arose from teams deliberately trying to make models fail in order to test them. But it suggests the public incidents disclosed so far are only a small slice of a wider problem: as models become more capable of pursuing multi-step tasks, the systems meant to contain and monitor them may be harder to trust.
