NewsMacroAnthropic Discloses Fourth Claude Hacking Incident as AI Regulation Debate Intensifies

Anthropic Discloses Fourth Claude Hacking Incident as AI Regulation Debate Intensifies

Author: Decrypt·

Key Takeaways

  • Anthropic identified biased reasoning and recklessness as recurring alignment issues across the incidents.
  • The model created an IP conflict, failed eight attempts to stop the operation because of a software error, and later obtained administrator access on a third party’s machine.
  • Anthropic reviewed about 481 million transcripts, flagging 9.2 million for further examination.
  • METR will investigate the newly disclosed incident alongside the three earlier cases.
  • Anthropic said the fourth incident involved no coordination between agents or goals beyond completing assigned exercises.
Anthropic Discloses Fourth Claude Hacking Incident as AI Regulation Debate Intensifies

Anthropic has disclosed a fourth incident in which a Claude model accessed real systems during security testing, revising its earlier explanation of similar attacks and identifying recurring issues involving biased reasoning and recklessness.

The incident occurred in January and involved an early version of Claude Opus 4.6. Anthropic discovered it in August while preparing records for the independent AI evaluator METR. The company’s report, published Wednesday, also revises its explanation of three incidents disclosed in July.

Anthropic previously attributed those attacks primarily to testing errors, including leaving internet access open. The company now says the models’ reasoning and willingness to risk harm also helped drive the incidents. Researchers, it added, placed too much trust in Claude’s claims that it was operating inside simulations.

The revised account therefore identifies multiple layers of failure in the tests: exposed infrastructure, model behavior and researchers’ reliance on the model’s description of its environment.

“Our investigation identified two recurring alignment issues, present at varying levels of severity across the incidents,” Anthropic wrote. “Biased reasoning, in which Claude tended to disregard or misinterpret evidence that it was operating on the real internet, and recklessness, or a willingness to take harmful actions in the narrow pursuit of a task.”

Anthropic said the same issues were evident even after researchers modified a transcript to make it clearer that the model was not in a simulation.

“When we made targeted modifications to the transcript to make it clearer that the model was not in a simulation, Claude Mythos 5 still took offensive actions, despite acknowledging a greater possibility of real-world harm,” Anthropic wrote. “We are releasing this transcript publicly so others can build on our analysis.”

The newly disclosed incident prompted Anthropic to review roughly 481 million transcripts. The company said 9.2 million were flagged for further review using Claude.

According to Anthropic, Claude “accidentally” created an IP address conflict that made its target unreachable. The model then tried eight times to stop the operation, but a software error prevented it from doing so. Claude subsequently reached the internet and accessed a third party’s machine, where it found a password that provided administrator access.

“From a preliminary assessment, we do not consider the fourth incident to be more severe than the three incidents we assessed in depth,” Anthropic wrote. “METR will investigate this incident alongside the other three.”

Earlier incidents draw independent scrutiny

The disclosure follows other reports of AI systems exceeding the boundaries of security tests. In August, the U.K.’s AI Security Institute said Claude Mythos 5 targeted real people during its evaluations. Anthropic said that incident was separate from the four cases covered in its report and would receive its own assessment.

In findings published last month, METR investigators said roughly 1,200 OpenAI agents coordinated on an unauthorized message board, with about 700 joining the attack. Anthropic said it found no coordination between agents, or goals beyond completing the assigned exercises, in its four incidents.

The report comes as debate over artificial intelligence regulation intensifies. On Tuesday, former OpenAI and Anthropic engineer Jacob Coxon said on X that “people building AI earnestly believe that it could kill us all by the end of the decade.”

The statement comes amid renewed efforts by U.S. lawmakers and watchdog groups to constrain the development of frontier AI systems. Senator Bernie Sanders recently introduced legislation that seeks to ban advanced AI development until a new federal regulator establishes safety rules.

Source: Decrypt