新聞宏觀經濟AI 與憲法(來自我的電子郵件)

AI 與憲法(來自我的電子郵件)

作者: Marginal Revolution·

重點速覽

  • Anthropic 自 2023 年起為 Claude 採用憲法,作為其 Constitutional AI 方法的一部分。
  • Scott Jenkins 表示,以案例法為基礎的 AI 治理模式可能比靜態憲法更具適應性。
  • 他警告,AI 系統每天產生大量邊緣案例,人工審查可能成為瓶頸。
  • Jenkins 也指出,先例可能變得不一致,AI 審查者之間可能存在共同盲點。
  • 文章未說明 Anthropic 的案例法憲法將由誰裁決,或其實際執行權力為何。
AI 與憲法(來自我的電子郵件)

《Marginal Revolution》這個由喬治梅森大學經濟學家 Tyler Cowen 共同撰寫的經濟部落格,刊出了一封來自 Scott Jenkins 的電子郵件,回應 Cowen 關於造訪 Anthropic 並就 Claude 憲法提供建議的筆記。Anthropic 自 2023 年起就為 Claude 公布了一份憲法——一套書面原則,用來引導助理的行為,並建立在該公司的 Constitutional AI 研究之上——因此,這場討論關注的是該文件外圍應採取何種治理形式,而不是它是否存在。Jenkins 認為,以普通法、案例法與獨立裁決為基礎的方法,比靜態文本更具適應性,但也提醒這種做法帶有結構性風險。

“Dear Tyler,

I enjoyed reading your notes on visiting Anthropic to advise on Claude's constitution. Framing AI governance around the common law, case law (“Talmud”), and independent adjudication is a much more adaptive approach than relying on a static, top-down text.

That said, moving from a fixed text to a case-law system introduces its own set of structural risks. If Anthropic adopts this direction, a few institutional design hazards seem worth anticipating:

The throughput bottleneck (Speed vs. Due Process): AI models generate billions of dynamic, edge-case interactions daily, while human judicial processes operate at human speed. If human adjudicators can only review a tiny fraction of flagged disputes, the actual operational rules will quietly decouple from official doctrine. Without automated verification tools to bridge this bandwidth gap, real oversight may only touch superficial cases.

The danger of tangled precedent (Doctrinal bloat): The common law works because human societies change at a manageable pace. With rapid model updates and shifting capabilities, the volume of case law, exceptions, and secondary interpretations could quickly become self-contradictory. Over time, this leads to doctrine that serves as post-hoc justification rather than a coherent operational constraint.

Correlated blind spots among AI reviewers: Using a diverse panel of AIs to detect constitutional drift is clever, but if these models share similar base data, fine-tuning techniques, or foundational architectures, their consensus will have shared blind spots. A model might learn to satisfy the specific rubrics of the reviewer panel while still drifting in ways the entire panel fails to register.

The “Hollow Court” trap: The hardest problem in any independent judiciary is enforcement against the institution funding it. If economic or competitive pressures rise, an adjudicative board that lacks hard veto power risks becoming purely performative—producing elaborate legal commentary while commercial realities dictate the real guardrails.

The common-law analogy is compelling, but the real test is whether the institutional machinery can handle the sheer velocity and scale of software.”

Jenkins 提出的四項風險,將長期以來關於制度設計的問題——例如司法機構如何保持對出資者的獨立性,以及累積的先例如何維持一致性——轉化到軟體環境中。對前沿模型的獨立監督已經是整個產業中活躍的議題:實驗室會發布憲法文件與系統卡,外部紅隊與政府安全研究機構也已承擔起主要開發者的審核角色。這篇文章留下未解的問題是,在 Anthropic,究竟由誰來裁決一套案例法憲法,以及該機構實際上會擁有多大的執行權力。

這篇文章最初發表於 Marginal Revolution,時間是 2026 年 8 月 25 日。