Anthropic 究竟为 AI 安全做了什么?What Exactly Is Anthropic Doing About AI Safety?
在 Anthropic 这里,安全不是一个部门,而是整个产品叙事。下面把宪法式 AI、红队测试与负责任扩展政策串起来看,也看看它们对「安全模型」的未来意味着什么。
At Anthropic, safety is not a department — it is the product story. Here is how Constitutional AI, red teaming and the Responsible Scaling Policy fit together, and what that means for the future of secure models.
这个领域的每家实验室都在追逐自己所能造出的最强系统,其中却有一家额外立了规矩:一旦事情变得危险,就必须自己踩刹车。这家公司就是 Anthropic——由一批出身 OpenAI 的人创办,并把「安全优先」当作自己在行业里的卖点。
我们先前那篇 AI 给普通人带来的风险 讲的是用户端能感知到的部分;Anthropic 处理的层级更靠下,动的是地基。但抛开口号,「安全优先」落到实操里究竟是什么样?下面完整走一遍:它的技术手法、内部政策,以及支撑 2026 年整套策略的理念。
01Anthropic 最初想做什么
公司于 2021 年成立,由 Dario 与 Daniela Amodei 领衔,团队里还有一批从 OpenAI 出来的研究者。所有工作都基于同一个判断:AI 系统会越来越强,而一种不服从任何人类价值的强大力量,可能造成灾难性的后果。
业内不少地方是等产品成型后再补安全,或者记者来问了才拿出来讲;在这里,安全直接摆在商业模式的正中央。公司的论点是:道德义务和好产品恰好是一回事——一个会编造事实、张口骂人、甚至帮人犯罪的软件,本身就是设计失败。关于这一使命的官方表述,可以看它的 使命说明页。
02宪法式 AI(CAI)是怎么回事
Anthropic 研究体系的核心成果就叫「宪法式 AI」。要理解它为什么重要,得先看常规的训练流程。我们曾在 各家实验室如何让模型变安全 里介绍过标准的 RLHF——把人类留在打分环节里。Anthropic 把这一步换成了 RLAIF,即「基于 AI 反馈的强化学习」。
- ❓
有害请求
→
🤖
模型起草回答
→
📜
对照规则手册
→
✅
自行改写
规则手册里写了什么
模型会拿到一份成文的原则清单,也就是它的「宪法」。每当生成一段回答,它还要按这些原则给这段回答打分。如果草稿违反了某条规则(最常被举的例子是「绝不解释炸弹怎么做」),模型就得改写,直到内容既无害又不失用处。这样一来,数百万小时的人工标注就被省掉了——模型在很大程度上是自己教自己。这份宪法的现行全文在 Anthropic 的 研究公告 里公开,任何人都可以免费阅读。
03把模型往死里测
很少有实验室把红队测试做得这么狠。在 Claude 这类系统发布之前,数百位专家——生物安全学者、网络安全黑客等等——被允许放手去攻破它。测试结果还会进入公司与政府评估机构之间签订的部署前安排。其中一家是 CAISI,即美国「人工智能标准与创新中心」,它设在 联邦标准机构——美国国家标准与技术研究院——职责是在模型发布前后对它进行审查。
生物安全测试
网络漏洞
社会工程
越狱攻击
04负责任扩展政策(RSP)
最能把 Anthropic 与其他公司区分开的,大概就是「负责任扩展政策」。这套框架由公司自行制定、完全出于自愿,2023 年首次发布,此后经过多轮修订。它的任务是:当先进系统越过一道道既定的能力门槛——也就是「AI 安全等级」(ASL)——时,控制由此带来的灾难性风险。最新完整文本可在 这一页 查看。
如果某个模型被证明能实质性地协助网络攻击、说服他人加入极端组织,或为生物武器提供帮助,这套政策就要求 Anthropic 在继续部署之前先强化防护措施。可以把它理解成一道安全刹车——而几乎所有竞争对手都拒绝正式装上一道。
05游说、监管与全球格局
内部的技术修补只是策略的一半,Anthropic 还投入相当精力去游说政府加强监管。它公开发表的政策文件主张:安全测试应当强制、最强的那批模型需要许可证,国家之间也要开展合作。
这一立场与当下各国陆续出台的框架高度契合,比如我们前不久拆解过的 欧盟 AI 法案通俗版。法规原文也可以直接从欧盟委员会的 官方条例文本 读取。在 Anthropic 看来,行业自律远远不够——要让 AI 朝着有利于全人类的方向发展,民主政府必须坐到牌桌上。
06最常被问到的几个问题
Anthropic 在 AI 安全上的主要赌注是什么?
「负责任扩展政策」到底承诺了什么?
宪法式 AI 与标准 RLHF 的区别在哪里?
Anthropic 的 AI 能算完全安全吗?
Anthropic 对政府监管持什么态度?
Every lab in this field is chasing the most capable system it can build. One of them has also written down rules for stopping itself when the work turns dangerous. That lab is Anthropic — launched by people who previously worked at OpenAI — and it markets itself as the industry's "safety-first" option.
Our earlier piece on the everyday risks AI poses covered what surfaces at the user's end. Anthropic works further down the stack, at the foundations. What does "safety-first" mean once you look past the slogan, though? Here is the full tour: the methods, the internal policies and the philosophy behind the company's 2026 strategy.
01What Anthropic Set Out to Do
The company opened its doors in 2021. Dario and Daniela Amodei led the effort, joined by other researchers who had come over from OpenAI. Everything rests on a single premise: AI systems will keep growing more powerful, and power that answers to no human values can do catastrophic damage.
In much of the industry safety gets bolted on late, or trotted out when a reporter calls. Here it sits at the centre of the business model. The company's argument is that the ethical obligation and the good product happen to be the same thing — software that fabricates facts, emits abuse or pitches in on a crime is broken by design. You can read the company's own account of this mission on its mission statement page.
02Constitutional AI (CAI), Explained
The centrepiece of Anthropic's research programme carries the name "Constitutional AI." To see why it matters, start with the usual training pipeline. Standard RLHF, the human-feedback route we described in how labs make their models safe, keeps people inside the ranking loop. Anthropic swaps that out for RLAIF — Reinforcement Learning from AI Feedback.
- ❓
Harmful Request
→
🤖
Model Drafts Reply
→
📜
Rulebook Check
→
✅
Rewrites Itself
Inside the Rulebook
The model is handed a written set of principles — its "constitution." Once an answer exists, the model is asked to grade that answer against those principles. A draft that breaks a rule (the standard illustration is "never explain how to build a bomb") gets revised until it is harmless yet still useful. Millions of hours of human labelling are thereby avoided; the model largely trains itself. Anthropic publishes the constitution's latest version at its research announcement — the whole text, free for anyone who cares to read it.
03Stress-Testing the Model
Few labs red-team as aggressively. Before a system such as Claude ships, hundreds of specialists — biosecurity experts, cybersecurity hackers and others — are handed licence to break it. Their findings also feed the pre-deployment arrangements the company holds with government evaluators. One such body is CAISI, the Center for AI Standards and Innovation, a U.S. agency housed at the federal standards body — the National Institute of Standards and Technology — whose remit is to inspect frontier-lab models both before and after release.
Biosecurity Tests
Cyber Vulnerabilities
Social Engineering
Jailbreaking
04The Responsible Scaling Policy, or RSP
The commitment that sets Anthropic apart may well be its "Responsible Scaling Policy." Written by the company itself and entirely voluntary, the framework first appeared in 2023 and has been revised repeatedly since. Its job is to manage catastrophic risk as advanced systems push past defined capability thresholds — the "AI Safety Level" (ASL) marks. The complete, up-to-date text is available at this page.
Should a model prove that it could meaningfully assist a cyberattack, talk people into joining an extremist group, or contribute to a bioweapon, the policy obliges Anthropic to strengthen its safeguards before deployment continues. Think of it as a safety brake — one that nearly every rival lab has declined to install formally.
05Lobbying, Regulation and the Global Picture
Internal technical fixes are only part of the strategy; Anthropic also spends real effort lobbying for government oversight. Its published policy papers call for mandatory safety testing, licences for the most capable models, and cooperation between countries.
That position lines up neatly with frameworks now emerging worldwide, such as the EU AI Act, put plainly, which we broke down recently. The regulation's own wording can be read straight from the European Commission's official regulation text. Self-regulation, in Anthropic's view, does not go far enough — democratic governments need a seat at the table if AI is to develop in humanity's favour.