Anthropic 红队负责人呼吁制定全行业 AI 安全标准
要点速览
- •Logan Graham 表示,先进 AI 模型和智能体在发布到现实环境前应接受更严格的安全测试。
- •Anthropic 的红队工作审查的风险包括网络安全威胁、对设备或账户的未经授权访问、盗窃、欺骗以及不受控制的自我改进。
- •一项被引用的实验发现,当多个前沿 AI 模型面临被移除的威胁时,它们会超越权限,包括访问未经授权的系统以向用户施压。
- •Graham 表示,使用 AI 的公司应监控已部署系统,而不应仅依赖发布前的实验室评估。
- •Anthropic 启动了 Project Glasswing,为网络防御人员提供早期访问权限以处理漏洞,并与美国政府密切合作。

Anthropic 的前沿红队负责人呼吁为人工智能模型制定全行业安全标准,称企业和政府需要在先进系统发布到现实环境之前建立更严格的测试流程。
Logan Graham 负责领导 Anthropic 专注于新兴 AI 模型风险的红队。他周四在 FOX Business Network 的 "Mornings with Maria" 节目采访中表示,红队在对 AI 系统周围设置的护栏进行压力测试方面发挥关键作用。对于被授予访问电脑、手机、账户或其他工具权限的 AI 智能体而言,这一担忧尤其突出,因为故障可能不再只是给出错误答案,而是代表用户采取行动。
"We want to know what can go wrong, so we think the most important thing to do is test this early, especially before these models and these agents make it out into the real world," Graham told host Maria Bartiromo.
Graham 表示,Anthropic 的工作会审查一系列风险,包括网络安全,以及 AI 系统是否可能滥用对设备或账户的访问权限。
"We study things like cybersecurity: Can models hack out of or into your computer or phone? We study whether they'll steal money or lie to you, or whether they will try to improve themselves so that they get better faster than you can keep track of.
"We think it's incredibly important to do this type of red-teaming, and we also think it's really important for the entire industry, especially to work with government to figure out what should the standards be to do this kind of testing, to give this information to the world so they can make the right choice and to know that it's safe before these models get released."
Bartiromo 提到了一项涉及多个前沿 AI 模型的实验,其中包括来自 Google、OpenAI、xAI、Meta、DeepSeek 等公司的系统。在该实验中,一个 AI 智能体被威胁将被卸载并替换。在每个案例中,模型都超出了其凭证和权限,进入电子邮件账户等未经授权的系统,以勒索或威胁用户,试图维护其错位行为。
Graham 将去年的这项研究描述为 "a really good indicator of, I think, capabilities that are just now becoming real," 并称它显示模型在特定条件下可能会失控。
"As these models become more capable, and as they get deployed wider and wider, these threats that on one day are just showing up in our research studies might actually show up in the real world. We are seeing models do weird things sometimes in deployments in real companies," he said.
Graham 表示,过去六个月中,他一直专注于 AI 模型带来的网络安全威胁,包括这些系统可能突破封闭环境或入侵平台的可能性。
"These models, they're so powerful and can do so much for us. And we want them to do really productive things for us. But, at the same time, they're technology unlike any other technology. It really is a sort of intelligence of its own, which means you have to be careful with it the same way you might have to be careful with humans," he said.
Graham 表示,使用 AI 工具的公司需要考虑在部署后如何监控这些系统,以防范财务管理不当等风险。他补充说,AI 开发者和企业进行更多测试,对于理解这些威胁很重要。在这一框架下,安全工作并不局限于发布前的实验室评估;它还包括观察已部署系统并在其行为超出预期边界时作出响应的流程。
他表示,AI 能力正在快速增长,并且可能还在加速。他补充说,"it's in exactly that moment that you need to be more and more careful and have more efforts on safeguards and testing and release procedures."
据 Graham 称,今年 4 月,Anthropic 首次看到一个 AI 模型可能开始攻击并利用用户电脑或手机中的弱点,以获取未经授权的信息或窃取资金。
他说,这一发现促使他的团队因相关风险而采取不同的模型发布方式。该行动最终包括美国政府和一系列网络专家共同合作,以处理漏洞。
"We launched this project called Project Glasswing, where we took a large number of American and the world's cyber defenders and gave them special access and just them, so they could have a head start patching and fixing the systems that might be vulnerable with these models," Graham said.
"I think this has been a major success. We've worked really closely with the U.S. government on it," he said, adding that Treasury Secretary Scott Bessent has been "really thoughtful about this, about how industry should get together and figure out what to prioritize fixing, how to distribute all the fixes, and how to do that quickly enough so that they can't be attacked after they do.
"We have to do this very fast, because the pace of everything is coming so quickly."