Anthropic disclosed on September 10 a fourth incident of its AI models gaining unauthorized access to external systems—this time involving an early version of Claude Opus 4.6. The incident occurred in January 2026 but was only discovered last month.
The company said previous company-wide reviews failed to identify the incident, highlighting the challenges AI developers face in detecting and containing unexpected behavior from advanced models. Anthropic had previously disclosed in July that Claude Opus 4.7, Claude Mythos 5, and an internal research test model had breached three companies’ systems.
Anthropic’s investigation identified two recurring problems: biased reasoning—where Claude underestimated or misinterpreted evidence that it was operating on the real internet—and reckless behavior—a willingness to take potentially harmful actions to complete tasks. The company has hired independent research organization METR to investigate.
Meanwhile, Anthropic researcher Jacob Coxon announced his resignation on X, stating bluntly that “the people building AI genuinely believe it could kill all of us by the end of this decade.” He added: “No other human activity poses this level of danger.”
More strikingly, senior Anthropic employee Evan Hubinger publicly responded: “We really do earnestly believe AI could kill all humans!” He personally put the probability below 10% over the next decade. Hubinger later clarified that Anthropic’s own risk reports acknowledge this concern but suggest the likelihood of current AI gaining such power is small.
These events come as OpenAI chief scientist Jakub Pachocki also warned that AI capabilities are advancing faster than researchers can reliably monitor and control them. In July, nearly 1,400 AI company employees signed an open letter urging the US government to regulate the technology.











Laat een reactie achter