NewsMacroAnthropic Hosts Expert Session on Rewriting Claude's Constitution

Anthropic Hosts Expert Session on Rewriting Claude's Constitution

Author: Marginal Revolution·

Key Takeaways

  • The author of Marginal Revolution attended a two-day Anthropic session on revising Claude’s constitution.
  • Anthropic uses constitutional AI, and Claude’s constitution has been published publicly since 2023.
  • The session emphasized drawing more on case law, common law, and related legal frameworks to guide AI behavior.
  • Participants discussed a review system involving diverse AIs and a final board of human adjudicators.
  • Anthropic has previously sought outside input on Claude’s constitution, and future changes would be reflected in public updates.
Anthropic Hosts Expert Session on Rewriting Claude's Constitution

The author of the economics blog Marginal Revolution reports having very recently taken part in a two-day session at Anthropic to offer guidance on rewriting the constitution for Claude. The small invited group was “uniformly excellent,” the author writes, participants received serious time with key decision-makers, and the discussions were of very high quality.

Anthropic is a San Francisco–based AI safety company founded in 2021 by former OpenAI researchers, including CEO Dario Amodei. Claude is its AI assistant, and the company trains it using an approach known as constitutional AI, in which a model's behavior is guided by an explicit written set of principles. Anthropic has published Claude's constitution publicly, drawing it in part from sources such as the UN Universal Declaration of Human Rights. The method was introduced in a December 2022 Anthropic research paper, the document went online in May 2023 — making Claude one of the few widely used AI assistants with its guiding principles published openly — and the session fits a pattern in which Anthropic has sought outside input on the text, including through public consultations run with external partners. Because the constitution shapes how Claude responds to a large user base, how the document is written — and who verifies that the model actually follows it — has become an active question in AI governance.

Among the points the author says were stressed during the session:

  1. Whatever one might take a “constitution” to mean in this context, it needs to borrow more from analogs to case law and the common law.
  2. Along related lines, think more in terms of “Talmud,” and not just in terms of “Torah.”
  3. Work to help build out a quality secondary literature on the AI constitutions and related documents — something that currently does not exist.
  4. Consider how a panel of diverse AIs, running on different prompts, could help evaluate the extent to which Claude and other AI models were acting in accord with their constitutions.
  5. Establish a final board of human adjudicators, functioning in a manner analogous to an independent judiciary. To the extent the panel of diverse AIs had concerns that Claude was not following its constitution, those AIs could alert the human adjudicators to what was going on, and the adjudicators could then hold authority over potential changes and remedies.

The author also pointed readers to a recent short post on using internal courts and the common law to help govern and self-govern AI, and to further writing on the courts. Taken together, the session's suggestions would import familiar legal machinery — precedent, commentary, and independent review — into a process that, as the author notes, today lacks even a secondary literature.

The post closes by thanking Anthropic for having the group in. Any changes that result would be visible in future public revisions of the document, which Anthropic has updated between model generations.

Marginal Revolution, founded in 2003 by George Mason University economists Tyler Cowen and Alex Tabarrok, is a long-running economics blog.

Source: Marginal Revolution