Anthropic 究竟为 AI 安全做了什么?What Exactly Is Anthropic Doing About AI Safety?

🛡️ AI 安全⏱11 分钟阅读📅更新于 2026 年 6 月

在 Anthropic 这里,安全不是一个部门,而是整个产品叙事。下面把宪法式 AI、红队测试与负责任扩展政策串起来看,也看看它们对「安全模型」的未来意味着什么。

◆知微•🛡️ AI 安全 · ⏱11 分钟阅读 · 2026 年 6 月 23 日
🛡️ AI Safety⏱ 11 min read📅 Updated June 2026

At Anthropic, safety is not a department — it is the product story. Here is how Constitutional AI, red teaming and the Responsible Scaling Policy fit together, and what that means for the future of secure models.

◆知微•🛡️ AI Safety · ⏱ 11 min read · June 23, 2026

这个领域的每家实验室都在追逐自己所能造出的最强系统,其中却有一家额外立了规矩:一旦事情变得危险,就必须自己踩刹车。这家公司就是 Anthropic——由一批出身 OpenAI 的人创办,并把「安全优先」当作自己在行业里的卖点。

我们先前那篇 AI 给普通人带来的风险 讲的是用户端能感知到的部分;Anthropic 处理的层级更靠下,动的是地基。但抛开口号,「安全优先」落到实操里究竟是什么样?下面完整走一遍:它的技术手法、内部政策,以及支撑 2026 年整套策略的理念。

01Anthropic 最初想做什么

公司于 2021 年成立,由 Dario 与 Daniela Amodei 领衔,团队里还有一批从 OpenAI 出来的研究者。所有工作都基于同一个判断:AI 系统会越来越强,而一种不服从任何人类价值的强大力量,可能造成灾难性的后果。

业内不少地方是等产品成型后再补安全,或者记者来问了才拿出来讲;在这里,安全直接摆在商业模式的正中央。公司的论点是:道德义务和好产品恰好是一回事——一个会编造事实、张口骂人、甚至帮人犯罪的软件,本身就是设计失败。关于这一使命的官方表述,可以看它的 使命说明页。

02宪法式 AI(CAI)是怎么回事

Anthropic 研究体系的核心成果就叫「宪法式 AI」。要理解它为什么重要,得先看常规的训练流程。我们曾在 各家实验室如何让模型变安全 里介绍过标准的 RLHF——把人类留在打分环节里。Anthropic 把这一步换成了 RLAIF,即「基于 AI 反馈的强化学习」。

宪法式 AI 背后的 RLAIF 循环
  1. ❓
    有害请求

    →

    🤖
    模型起草回答

    →

    📜
    对照规则手册

    →

    ✅
    自行改写

规则手册里写了什么

模型会拿到一份成文的原则清单,也就是它的「宪法」。每当生成一段回答,它还要按这些原则给这段回答打分。如果草稿违反了某条规则(最常被举的例子是「绝不解释炸弹怎么做」),模型就得改写,直到内容既无害又不失用处。这样一来,数百万小时的人工标注就被省掉了——模型在很大程度上是自己教自己。这份宪法的现行全文在 Anthropic 的 研究公告 里公开,任何人都可以免费阅读。

03把模型往死里测

很少有实验室把红队测试做得这么狠。在 Claude 这类系统发布之前,数百位专家——生物安全学者、网络安全黑客等等——被允许放手去攻破它。测试结果还会进入公司与政府评估机构之间签订的部署前安排。其中一家是 CAISI,即美国「人工智能标准与创新中心」,它设在 联邦标准机构——美国国家标准与技术研究院——职责是在模型发布前后对它进行审查。

🧬严重

生物安全测试

测试者会想办法让模型交出病原体或危险化学品的制作配方。
💻严重

网络漏洞

它能不能写出可用的恶意软件、把零日漏洞变成武器,或者给入侵行动搭把手?
🎭高

社会工程

这一项瞄准的是极具说服力的钓鱼邮件,以及用来操控舆论的政治宣传。
🧩高

越狱攻击

研究者用复杂的角色扮演和逻辑谜题,试图绕过守护模型的安全过滤器,以及它们背后的系统提示。

04负责任扩展政策(RSP)

最能把 Anthropic 与其他公司区分开的,大概就是「负责任扩展政策」。这套框架由公司自行制定、完全出于自愿,2023 年首次发布,此后经过多轮修订。它的任务是:当先进系统越过一道道既定的能力门槛——也就是「AI 安全等级」(ASL)——时,控制由此带来的灾难性风险。最新完整文本可在 这一页 查看。

5
AI 安全等级(ASL)
100%
承诺暂停
ASL-4
当前重点领域

如果某个模型被证明能实质性地协助网络攻击、说服他人加入极端组织,或为生物武器提供帮助,这套政策就要求 Anthropic 在继续部署之前先强化防护措施。可以把它理解成一道安全刹车——而几乎所有竞争对手都拒绝正式装上一道。

05游说、监管与全球格局

内部的技术修补只是策略的一半,Anthropic 还投入相当精力去游说政府加强监管。它公开发表的政策文件主张:安全测试应当强制、最强的那批模型需要许可证,国家之间也要开展合作。

这一立场与当下各国陆续出台的框架高度契合,比如我们前不久拆解过的 欧盟 AI 法案通俗版。法规原文也可以直接从欧盟委员会的 官方条例文本 读取。在 Anthropic 看来,行业自律远远不够——要让 AI 朝着有利于全人类的方向发展,民主政府必须坐到牌桌上。

06最常被问到的几个问题

Anthropic 在 AI 安全上的主要赌注是什么?
核心手段是「宪法式 AI」:用 AI 自己生成的反馈(RLAIF)训练模型遵守成文的伦理原则(即那部「宪法」),而不是只靠人工标注。在这之外,还有一套严格的负责任扩展政策兜底。
「负责任扩展政策」到底承诺了什么?
RSP 是一项自愿承诺:如果模型的能力达到某个会给公共安全带来严重风险的水平——比如协助网络攻击,或者帮助制造生物武器——Anthropic 承诺升级防护措施,必要时暂停开发。
宪法式 AI 与标准 RLHF 的区别在哪里?
标准 RLHF 由人来给 AI 的输出排序,宪法式 AI 则把这个环节交给 AI 自己(即 RLAIF)。模型拿自己写出的内容去对照一部规则「宪法」,于是对齐过程更快、更容易扩展,对人力劳动的依赖也更少。
Anthropic 的 AI 能算完全安全吗?
没有什么是 100% 安全的。Anthropic 在安全研究上处于领先,红队演练也相当严格,但它的模型依然会出错、会编造事实,也可能被复杂的越狱提示词操控。用户的警惕一刻都不能少。
Anthropic 对政府监管持什么态度?
是的。这家公司呼吁政府加强监督,主张强制安全测试,并要求对能力最强的模型实行许可制。在它看来,靠科技公司自律不足以管住先进 AI 带来的风险。
◆

知微

我们研究 AI 技术与企业安全政策,并为普通用户提供可操作的建议。本指南的准确性已于 2026 年 6 月复核。有疑问,或者想参与贡献?今天就 联系我们。

Every lab in this field is chasing the most capable system it can build. One of them has also written down rules for stopping itself when the work turns dangerous. That lab is Anthropic — launched by people who previously worked at OpenAI — and it markets itself as the industry's "safety-first" option.

Our earlier piece on the everyday risks AI poses covered what surfaces at the user's end. Anthropic works further down the stack, at the foundations. What does "safety-first" mean once you look past the slogan, though? Here is the full tour: the methods, the internal policies and the philosophy behind the company's 2026 strategy.

01What Anthropic Set Out to Do

The company opened its doors in 2021. Dario and Daniela Amodei led the effort, joined by other researchers who had come over from OpenAI. Everything rests on a single premise: AI systems will keep growing more powerful, and power that answers to no human values can do catastrophic damage.

In much of the industry safety gets bolted on late, or trotted out when a reporter calls. Here it sits at the centre of the business model. The company's argument is that the ethical obligation and the good product happen to be the same thing — software that fabricates facts, emits abuse or pitches in on a crime is broken by design. You can read the company's own account of this mission on its mission statement page.

02Constitutional AI (CAI), Explained

The centrepiece of Anthropic's research programme carries the name "Constitutional AI." To see why it matters, start with the usual training pipeline. Standard RLHF, the human-feedback route we described in how labs make their models safe, keeps people inside the ranking loop. Anthropic swaps that out for RLAIF — Reinforcement Learning from AI Feedback.

The RLAIF loop behind Constitutional AI
  1. ❓
    Harmful Request

    →

    🤖
    Model Drafts Reply

    →

    📜
    Rulebook Check

    →

    ✅
    Rewrites Itself

Inside the Rulebook

The model is handed a written set of principles — its "constitution." Once an answer exists, the model is asked to grade that answer against those principles. A draft that breaks a rule (the standard illustration is "never explain how to build a bomb") gets revised until it is harmless yet still useful. Millions of hours of human labelling are thereby avoided; the model largely trains itself. Anthropic publishes the constitution's latest version at its research announcement — the whole text, free for anyone who cares to read it.

03Stress-Testing the Model

Few labs red-team as aggressively. Before a system such as Claude ships, hundreds of specialists — biosecurity experts, cybersecurity hackers and others — are handed licence to break it. Their findings also feed the pre-deployment arrangements the company holds with government evaluators. One such body is CAISI, the Center for AI Standards and Innovation, a U.S. agency housed at the federal standards body — the National Institute of Standards and Technology — whose remit is to inspect frontier-lab models both before and after release.

🧬Critical

Biosecurity Tests

Testers push the model toward handing over recipes for pathogens or dangerous chemicals.
💻Critical

Cyber Vulnerabilities

Can it produce working malware, weaponise a zero-day flaw, or lend a hand to a break-in?
🎭High

Social Engineering

Here the aim is persuasive phishing mail, or political propaganda built to manipulate.
🧩High

Jailbreaking

Elaborate roleplay scenarios and logic puzzles are used to slip past the safety filters guarding a model, and the system prompts that sit behind them.

04The Responsible Scaling Policy, or RSP

The commitment that sets Anthropic apart may well be its "Responsible Scaling Policy." Written by the company itself and entirely voluntary, the framework first appeared in 2023 and has been revised repeatedly since. Its job is to manage catastrophic risk as advanced systems push past defined capability thresholds — the "AI Safety Level" (ASL) marks. The complete, up-to-date text is available at this page.

5
ASL (AI Safety Levels)
100%
Commitment to pause
ASL-4
Current focus area

Should a model prove that it could meaningfully assist a cyberattack, talk people into joining an extremist group, or contribute to a bioweapon, the policy obliges Anthropic to strengthen its safeguards before deployment continues. Think of it as a safety brake — one that nearly every rival lab has declined to install formally.

05Lobbying, Regulation and the Global Picture

Internal technical fixes are only part of the strategy; Anthropic also spends real effort lobbying for government oversight. Its published policy papers call for mandatory safety testing, licences for the most capable models, and cooperation between countries.

That position lines up neatly with frameworks now emerging worldwide, such as the EU AI Act, put plainly, which we broke down recently. The regulation's own wording can be read straight from the European Commission's official regulation text. Self-regulation, in Anthropic's view, does not go far enough — democratic governments need a seat at the table if AI is to develop in humanity's favour.

06Questions People Ask Most

Where does Anthropic place its main safety bet?
The core of the approach is "Constitutional AI," which trains a model to follow written ethical principles (a constitution) using feedback generated by AI (RLAIF) rather than depending on human labelling alone. A strict Responsible Scaling Policy backs it up.
And what does the Responsible Scaling Policy actually commit them to?
The RSP is a voluntary pledge: should a model reach a capability level that puts public safety at severe risk — assisting a cyberattack, say, or helping build a bioweapon — Anthropic promises to upgrade its safeguards and, if needed, pause development.
In what way does Constitutional AI depart from standard RLHF?
Standard RLHF has humans rank what the AI produces; Constitutional AI lets the AI do the ranking (RLAIF). The model measures its own output against a "constitution" of rules, which makes alignment quicker, easier to scale, and less dependent on human labour.
Can we call any Anthropic model completely safe?
Nothing is 100% safe. Anthropic leads the field in safety research and runs demanding red-team exercises, yet its models still slip up, invent facts, or get manipulated by elaborate jailbreaking prompts. Vigilance on the user's part never stops being necessary.
Where does Anthropic stand on government regulation?
Yes. Oversight from government is something the company pushes for, alongside compulsory safety testing and a licensing regime for the most capable systems. Its position is that self-regulation by tech firms leaves the risks of advanced AI insufficiently contained.
◆

知微

Our team digs into AI technology and the safety policies companies write, then turns that work into practical advice for ordinary users. Accuracy was checked in June 2026. Questions, or something to add? Get in touch with us today.