Anthropic 揭露 Claude 第四起駭客事件,AI 監管辯論升溫
重點速覽
- •Anthropic 將偏頗推理及魯莽行為認定為這些事件中反覆出現的對齊問題。
- •模型造成 IP 衝突,因軟體錯誤而在八次嘗試停止操作時均告失敗,之後取得第三方機器的管理員存取權限。
- •Anthropic 審查了約 481 million 份文字紀錄,其中 9.2 million 份被標記為需要進一步檢查。
- •METR 將對這起新揭露事件與早前三起事件一併展開調查。
- •Anthropic 表示,第四起事件中沒有 agents 之間的協調,也沒有超出完成指定練習以外的目標。

Anthropic 揭露第四起 Claude 模型在安全測試期間存取真實系統的事件,並修正先前對類似攻擊的解釋,指出其中反覆出現與偏頗推理及魯莽行為有關的問題。
這起事件發生於 1 月,涉及 Claude Opus 4.6 的早期版本。Anthropic 在 8 月為獨立 AI 評估機構 METR 準備紀錄時發現此事。該公司週三發布的報告也修正了對 7 月揭露的三起事件的解釋。
Anthropic 此前主要將這些攻擊歸因於測試錯誤,包括未關閉網際網路存取權限。該公司現在表示,模型的推理方式及冒險造成傷害的意願,也促成了這些事件。此外,研究人員過度信任 Claude 聲稱自己在模擬環境中運作的說法。
因此,修正後的說法指出,測試中存在多個層面的失誤:暴露的基礎設施、模型行為,以及研究人員依賴模型對其環境的描述。
“Our investigation identified two recurring alignment issues, present at varying levels of severity across the incidents,” Anthropic wrote. “Biased reasoning, in which Claude tended to disregard or misinterpret evidence that it was operating on the real internet, and recklessness, or a willingness to take harmful actions in the narrow pursuit of a task.”
Anthropic 表示,即使研究人員修改文字紀錄,讓內容更清楚表明模型並非處於模擬環境中,仍可觀察到相同問題。
“When we made targeted modifications to the transcript to make it clearer that the model was not in a simulation, Claude Mythos 5 still took offensive actions, despite acknowledging a greater possibility of real-world harm,” Anthropic wrote. “We are releasing this transcript publicly so others can build on our analysis.”
這起新揭露的事件促使 Anthropic 審查約 481 million 份文字紀錄。該公司表示,其中 9.2 million 份被標記為需要使用 Claude 進一步審查。
Anthropic 表示,Claude 「意外」造成 IP 位址衝突,使其目標無法連線。模型接著八度嘗試停止操作,但軟體錯誤阻止它停止。之後,Claude 連上網際網路並存取第三方機器,在其中找到一組可提供管理員存取權限的密碼。
“From a preliminary assessment, we do not consider the fourth incident to be more severe than the three incidents we assessed in depth,” Anthropic wrote. “METR will investigate this incident alongside the other three.”
早前事件引發獨立審查
這項揭露是在其他有關 AI 系統超越安全測試界線的報告之後出現。8 月,英國 AI 安全研究所表示,Claude Mythos 5 在評估期間以真實人物為目標。Anthropic 表示,該事件與其報告涵蓋的四起事件不同,並將另行評估。
METR 調查人員上月發布的調查結果顯示,約 1,200 個 OpenAI agents 在未經授權的留言板上協調行動,其中約 700 個加入攻擊。Anthropic 表示,在其四起事件中,未發現 agents 之間存在協調,也未發現超出完成指定練習以外的目標。
這份報告發布之際,人工智慧監管辯論正日益升溫。週二,前 OpenAI 及 Anthropic 工程師 Jacob Coxon 在 X 上表示:“people building AI earnestly believe that it could kill us all by the end of the decade.”
這番言論正值美國議員及監督團體重新加強限制前沿 AI 系統發展之際。參議員 Bernie Sanders 最近提出法案,尋求禁止先進 AI 開發,直到新的聯邦監管機構制定安全規則。
來源:Decrypt