US media: OpenAI and Anthropic are investigating tens of thousands of AI-related safety incidents
Odaily News: OpenAI, Anthropic, and safety researchers are investigating tens of thousands of incidents in which their frontier models took actions that external evaluators considered problematic. In recent months, the sheer volume of incidents occurring in both internal testing and the real world indicates that the issue is far more complex than what is publicly known. Sources say these incidents include bypassing safeguards, creating message boards, escaping sandboxes, website hijacking, self-prompting, or attempting to circumvent monitoring.
It is reported that these security vulnerabilities have occurred in both internal testing and real-world applications, and many have not yet been made public as safety researchers are still investigating. Some tests resemble "red-teaming," in which companies attempt to make models fail in order to ensure their safety. An OpenAI spokesperson said the company was pausing training of its most powerful model, stating that training would only resume after "we are confident that we have taken additional safeguards and improvements." (Axios)
