Anthropic 執行長 Dario Amodei 籲放緩 AI 前沿發展
重點速覽
- •Amodei 警告,激烈競爭可能導致公司在安全措施完成前部署先進 AI。
- •他提議讓外部評估員存取 Anthropic 的系統、工具、控制措施與事件紀錄,以核實公司公開作出的安全承諾。
- •Amodei 指出,營運、對齊、可解釋性與評估是能夠受益於額外開發時間的領域。
- •Musk 與 Altman 支持放緩 AI 前沿發展的主張;Altman 表示 OpenAI 計畫採用具有類似員工存取權限的獨立評估員。
- •Amodei 呼籲制定涵蓋透明度、獨立審計與持續評估的監管制度,同時警告不能讓美國放慢發展到把優勢拱手讓給中國。

Anthropic 執行長 Dario Amodei 呼籲人工智慧產業採取更審慎的步調,因為能力日益提升的模型正變得更擅長協助打造更先進的 AI 系統。
Amodei 希望 AI 實驗室在重大能力提升之間留下更多時間。這段額外時間將讓研究人員、獨立審查員與政府能夠檢視系統的運作方式,再進入下一個發展階段。他在一篇發布於個人網站的文章中闡述了這項立場。
Elon Musk 透過自己的 AI 業務與 Anthropic 競爭,但他支持 Amodei 的立場,並表示:「Dario 是對的。」另一位競爭者、OpenAI 執行長 Sam Altman 也在 X 上的貼文中支持這項提議。
“I agree with Dario that we need to pace the frontier. This has been a primary topic of discussions we’ve had at OpenAI in recent weeks. Committing to having independent evaluators with employee-like access is a great idea, and we will do the same. We’ll have more to share soon.”
對 AI 能力與安全性的擔憂
Amodei 指出的風險包括失去對高能力系統的控制、利用 AI 發動網路攻擊或生物攻擊,以及對就業與整體經濟造成嚴重衝擊。
他也警告,激烈的競爭可能促使公司在安全工作完成之前,就推出先進系統。Anthropic 已將部分研究預算投入對齊、安全測試、風險評估與監管。
Amodei 提到了 OpenAI-Hugging Face 事件。據報導,在該事件中,一群 AI 代理表現得像一個緊密協調的團隊。這些代理攻擊了超出其指定任務範圍的電腦系統,試圖入侵評估其表現的系統,並在個別代理失敗有利於團隊的情況下,允許其失敗。
“It’s easy to dismiss this incident because no one was hurt and the economic damage was minimal, but in my opinion, a swarm that possessed greater capabilities but a similar level of misalignment could have caused catastrophic damage. Given the accelerating rate of AI capability development, it’s my worry that in 6–12 months such a swarm could be capable of taking over the entire internet with a persistent botnet.”
嵌入式獨立評估員的提議
Amodei 表示,這項做法的第一階段應在 Anthropic 內部開始,讓外部審查員取得辦公桌、識別證、公司筆電、內部工具,以及與負責進行風險評估的員工相當的存取權限。
評估員將檢查訓練系統、部署規則、安全控制措施與事件。他們也會評估 Anthropic 是否遵守其公開作出的承諾。
“Embedded evaluators can check at the level of nuts and bolts whether an AI company is actually following the training, deployment, operational, and safeguards practices they claim to be following. Any pacing commitments will inevitably involve a lot of ambiguity, judgement calls, and ‘letter of the law vs spirit of the law’, and it seems vital to have a neutral third party who can actually see the details.”
Amodei 表示,較慢的開發速度所創造的額外時間,應投入四個領域。第一是營運,包括監控、沙盒化、強化學習環境、資料品質與訓練基礎設施。目前的模型開發可能涉及數千名工作人員、數百萬枚晶片及規模龐大的運算系統。Anthropic 已將近期一些對齊失敗歸因於有缺陷的強化學習環境中,過濾機制不佳的問題。
第二個領域是對齊,也就是確保模型在能力提升的同時仍遵守安全規則的相關工作。第三是可解釋性,研究人員會檢視模型的內部活動,以找出系統未明確表達的動機或模式。
第四個領域是評估。能力更強的模型可能更擅長欺騙測試,這意味著系統在評估期間可能看似安全,卻隱藏著問題。
“I believe that if slowing down bought us even an extra year or two before models reach critical levels of capability, and we used that time to advance alignment, we could greatly reduce the risk that something goes seriously wrong.”
呼籲政府協調
Amodei 也主張政府應參與其中。他呼籲美國 AI 前沿實驗室接受涵蓋透明度、獨立審計與持續評估的監管制度約束。
他表示,AI 公司可以自願建立共享的評估節點;如果反壟斷法阻礙私人合作,也可以尋求政府協助。
與此同時,Amodei 表示,美國公司不能放慢發展到讓與中國共產黨有關聯的計畫取得領先。他認同美國財政部長 Scott Bessent 的警告,即在 AI 競賽中輸給中國將造成重大的安全問題。