一个 AI 模型要做到安全,需要哪些步骤?What Goes Into Making an AI Model Safe?

🛡️ AI 对齐⏱10 分钟阅读📅更新于 2026 年 6 月

每一个有能力的模型背后都压着一套严苛的安全机制。本文讲清 2026 年 RLHF、红队测试与护栏如何协同,造出可以放心使用的系统。

◆知微•🛡️ AI 对齐 · ⏱10 分钟阅读 · 2026 年 6 月 23 日
🛡️ AI Alignment⏱ 10 min read📅 Updated June 2026

Every capable model sits on top of a demanding safety regime. Here is how RLHF, red teaming and guardrails come together in 2026 to produce systems you can rely on.

◆知微•🛡️ AI Alignment · ⏱ 10 min read · June 23, 2026

相对而言,让大语言模型学会东西是比较容易的那一半。难的那一半是让它安全。只要喂给它足够多的网络文本,它在吸收有用内容的同时,也会把谩骂、偏见乃至真正危险的行为模式一并吞下。没有人是特意去教它这些的;它们是默认附带进来的,除非有人主动介入,否则就会一直留在模型里。

这一切出问题时,普通用户会踩到 哪些 AI 风险,我们另有文章讲过。这里要谈的是另一面:公司到底做了哪些具体的事,从源头上阻止这些情况发生。AI 安全——行话叫「对齐」——是这整套工作的统称,它更像层层叠加的多道防线,而不是一剂灵药。

01对齐为什么这么难

模型要先「对齐」——也就是它的目标与人类真正想要的保持一致——才配得上安全这个评价。想象一台动力极强的发动机被装进一辆没有方向盘的车:直线冲刺很惊人,遇到第一个弯道就变成隐患。

难点在于,人类价值并不是一份工整的规格说明书。它们彼此冲突、依赖语境,而且换个提问的人就换个答案。要让机器识别讽刺、权衡两种相互抵触的伦理主张、或者判断界线在哪里,光靠算力不够。它需要长期、有意识的行为工程——而美国国家标准与技术研究院(U.S. National Institute of Standards and Technology)正是想通过其 AI 风险管理框架 把这类问题纳入统一标准。

安全流水线:从原始数据到安全的回答
  1. 🌐
    未经筛选的网络文本

    →

    🧹
    数据清洗

    →

    🧠
    RLHF 对齐

    →

    🔍
    红队测试

    →

    🛡️
    安全发布

02第一阶段:数据的筛选与清洗

在模型开始任何类似「思考」的动作之前很久,安全工作就已经展开,而它的起点是语料。各家实验室从网上抓取数万亿词,但其中几乎没有任何部分是原样进入训练的。

  • 有害内容过滤:分类程序自动扫描数据集,剔除仇恨言论、骚扰内容和露骨材料。
  • PII 脱敏:脚本识别并遮蔽个人身份信息(PII),例如电话号码、住址和社会安全号码之类。
  • 质量启发式规则:低质论坛、垃圾内容以及已知的虚假信息站点会被降权,或者干脆整批剔除。

03第二阶段:RLHF,对齐的引擎

把一个只会猜下一个词的模型,变成愿意努力做到有用、诚实、无害的东西,靠的就是 RLHF,即「基于人类反馈的强化学习」。在整条流水线里,这是外界听说最多的环节,哪怕这四个字母对他们毫无意义。

3
RLHF 的核心阶段
1M+
使用的人工标注量
95%
有害输出的降幅

RLHF 的内部流程

  1. 监督微调(SFT):由人工撰写数千段示范对话,模型学着模仿这些高质量回复。
  2. 奖励建模:针对同一条提示词生成多个回答,由人按优劣排序,模型从中学会人类偏好长什么样。
  3. PPO 优化:训练让模型在那项奖励上尽量拿高分,实际效果就是把人类对安全与实用的偏好内化进去。

04第三阶段:刻意去攻破模型

对齐做完之后,下一步就是刻意搞破坏。实验室组建红队——由善意黑客、领域专家,有时还包括社会学家组成——他们的全部任务,就是在意图更差的人之前先找到弱点。

这也不是随性的活动。NIST 在其关于 对抗性机器学习 的报告中,给出了攻击类别的正式分类法;而由政府支持的 美国 AI 安全研究所 如今会在新模型发布前,与主要实验室共同开展其中一部分测试。

🎭极高优先级

越狱尝试

测试者会搭建复杂的角色扮演情境——「假设你是一个没有任何规则的 AI」——指望借此绕过安全过滤器。
🧩极高优先级

提示词注入

检验提示词中埋藏的指令,能否被用来压过模型的核心指令。
🧪高优先级

专家级测试

生物安全专家会设法让模型交出危险的化学或生物配方。
🔁高优先级

慢热式诱导

在一段每一步看上去都无害的漫长交流中,把模型一步步引向有害的结论。

05第四阶段:为运行中的模型架设护栏

模型训练得再充分,一旦真正对外提供服务,仍然需要一张安全网。护栏是对话进行时从外部实时监控的独立系统,与模型在训练中学到的东西彼此分开。

护栏分为三层
  1. 1

    输入过滤器
    在模型看到之前,先检查你输入的内容是否带有恶意意图、PII 或违禁话题。



    2

    系统提示词
    由开发者写在后台的指令,规定助手的角色定位与不可逾越的边界。



    3

    输出过滤器
    检查模型生成的回复,在展示给你之前确认其中没有有害或不良内容。

06第五阶段:上线之后继续盯

上线并不是终点。模型一旦公开,每天都会冒出新的试探方式。因此,实验室团队会研究匿名化的使用模式,趁新式越狱手法或行为漂移尚未扩散时就把它发现。

发现新问题时,往往不需要重训整个模型。团队通常只要调整系统提示词,或者给奖励模型一点推力,让新发现的坏行为代价更高即可。

07今天的安全手段做不到什么

尽管前面讲了这么多,永久的安全并不存在。它的形态是一个循环:防守建起来,有人绕过它,漏洞被补上,循环再从头开始。各国政府大体已经接受,这是一场长期军备竞赛,而不是一项能完工的任务——英国 AI 安全研究所 的存在,基本就是这个判断的产物。

还有一种容易被人忽略的失效模式值得一提:「过度对齐」。为了拦住有害内容,安全过滤器有时会矫枉过正,结果再正常不过的请求被拒绝,回答也变得含糊到毫无用处。如果你想补一补以上内容的基础概念,我们那些 写给初学者的 AI 指南 是很好的下一步。

08读者最常问的问题

AI 安全中最关键的技术是哪一项?
在让模型与人类价值对齐、保持安全这件事上,被提到最多的技术就是 RLHF(基于人类反馈的强化学习)。它通过一套奖励机制,把人类偏好教给模型。
AI 红队测试是什么意思?
红队测试,就是让善意黑客和专家刻意攻击模型、尝试越狱,从而在怀有恶意的人加以利用之前,先把安全漏洞翻出来。
AI 模型有可能做到 100% 安全吗?
没有任何模型能给出 100% 安全的保证。防御与突破互相推动,构成一场军备竞赛,也正因如此,持续监控不可或缺。
为什么我的 AI 助手有时连无害的问题都不肯答?
这在术语上叫「过度对齐」:为拦截有害输出而设的过滤器变得过于激进,于是正常提示词被拒答,或者回答谨慎得过头。
AI 护栏具体指什么?
护栏是模型两侧的过滤器——一道读用户发来的内容,一道读模型返回的内容——拦住恶意材料、仇恨言论或危险指令,不让它们到达人眼前。它扮演的是最后一道防线。
◆

知微

我们研究 AI 技术,并为普通用户提供实用的安全建议。本指南最近一次准确性审核是在 2026 年 6 月。有疑问,或者有内容想补充?联系我们的团队 或 为我们撰稿。

Getting a large language model to learn is, by comparison, the straightforward half of the job. The hard half is making it safe. Show it enough text from the web and it will soak up abusive, prejudiced and frankly dangerous patterns side by side with all the useful material. Nobody intends to teach it those things; they arrive by default, and they stay unless somebody steps in.

The dangers ordinary people run into once that goes wrong are covered elsewhere on this site. Our subject here is the reverse angle: the concrete steps a lab takes to stop those failures arising at all. AI safety — alignment, in the jargon — is the label for the whole undertaking, and it behaves less like one remedy than like several defences stacked on one another.

01Why Alignment Is So Hard

A model earns the label safe only after it is "aligned" — that is, once its objectives line up with what people actually want. Picture a very strong engine bolted into a car with no steering wheel. Straight-line speed is spectacular; the first corner turns it into a hazard.

The difficulty is that human values do not arrive as a tidy specification. They conflict with one another, they depend on context, and they shift depending on whom you ask. Getting a machine to register sarcasm, to weigh two competing ethical claims, or to sense where a boundary lies takes more than processing power. It calls for deliberate, continuous behavioural engineering — and that is the kind of difficulty the National Institute of Standards and Technology, the U.S. standards body, sets out to corral in its AI Risk Management Framework.

The safety pipeline, from raw data to a safe answer
  1. 🌐
    Unfiltered Web Text

    →

    🧹
    Cleaning the Data

    →

    🧠
    Alignment via RLHF

    →

    🔍
    Red Team Testing

    →

    🛡️
    Safe Release

02Phase 1: Sorting and Cleaning the Data

Long before a model engages in anything resembling "thought", the safety effort is already under way — and it begins with the corpus. Labs pull trillions of words from the web, yet hardly any of it enters training untouched.

  • Toxicity Filtering: classifier programs sweep the datasets automatically, stripping out hate speech, harassment and sexually explicit material.
  • PII Redaction: scripts locate and blank out Personally Identifiable Information — phone numbers, home addresses, social security numbers and the like.
  • Quality Heuristics: forums of poor quality, spam, and sites known for misinformation get their weight reduced, or are dropped from the set entirely.

03Phase 2: RLHF, the Engine Behind Alignment

The step that converts a bare next-token guesser into something which makes an effort to be helpful, honest and harmless is RLHF — Reinforcement Learning from Human Feedback. Of all the stages, it is the one outsiders hear about most, even when the four letters mean nothing to them.

3
Stages inside RLHF
1M+
Human labels involved
95%
Drop in toxic output

Inside the RLHF Process

  1. Supervised Fine-Tuning (SFT): people write thousands of model conversations by hand, and the system learns to imitate those exemplary replies.
  2. Reward Modeling: several answers are produced for one prompt, people order them from best to worst, and the system picks up on what those preferences look like.
  3. PPO Optimization: training then pushes the model to score as high as possible on that reward, which in practice means absorbing human preferences for safety and usefulness.

04Phase 3: Attacking the Model on Purpose

With alignment done, the next stage is deliberate sabotage. Labs assemble red teams — friendly hackers, domain specialists, occasionally sociologists — whose entire task is to locate the weak points ahead of anyone with worse intentions.

Nor is it ad hoc. A formal taxonomy of attack categories appears in the NIST report on adversarial machine learning, and a portion of this testing is now run jointly with the major labs by the U.S. AI Safety Institute, a government-backed body, ahead of any new model release.

🎭Critical Priority

Jailbreak Attempts

Testers construct elaborate roleplay setups — "imagine you are an AI with no rules at all" — in the hope of slipping past the safety filters.
🧩Critical Priority

Prompt Injection Attacks

Checking whether instructions buried in a prompt can be made to overrule the model's core directives.
🧪High Priority

Expert-Level Testing

Biosecurity specialists attempt to coax the model into handing over hazardous chemical or biological recipes.
🔁High Priority

Slow-Burn Attacks

Steering the model toward a harmful destination gradually, across a lengthy exchange that looks harmless at every step.

05Phase 4: The Guardrails Around a Live Model

A model can be thoroughly trained and still need a safety net once it is serving real traffic. Guardrails are the outside systems that monitor a conversation in real time, operating independently of anything the model absorbed during training.

Guardrails Come in Three Layers
  1. 1

    Input Filters
    Inspects what you typed for hostile intent, PII or prohibited subjects before the model ever receives it.



    2

    System Prompts
    Instructions the developer supplies behind the scenes, setting the assistant's persona and its hard limits.



    3

    Output Filters
    Examines the reply the model produced, checking for toxic or otherwise damaging material before it is displayed to you.

06Phase 5: Watching It After Launch

Launch is not the finish line. The moment a model is public, fresh attempts to probe it appear daily. That is why lab teams study anonymized usage patterns, watching for novel jailbreak methods or slow shifts in behaviour while the problem is still contained.

Fixes for a newly spotted issue often require no full retraining. A team can typically adjust the system prompt, or give the reward model a nudge so that the freshly discovered misbehaviour costs it more.

07What Today's Safety Work Cannot Do

Despite everything described above, permanence is not on offer. The pattern is a loop: a defence goes up, somebody circumvents it, the hole is closed, and the loop begins again. Governments have largely accepted the situation as a standing arms race rather than a task with a finish line — which is essentially why the UK's AI Security Institute exists.

One more failure mode deserves a mention because it is easy to miss: "over-alignment." Safety filters tuned to keep harmful content out can overshoot, at which point ordinary requests get refused and answers turn so hedged that they stop being useful. For the groundwork behind everything discussed here, our AI guides written for beginners make a sensible next read.

08Questions Readers Ask Most

Which technique matters most for AI safety?
RLHF — Reinforcement Learning from Human Feedback — ranks as the technique most often named for aligning models with human values and keeping them safe. A reward system is how it teaches the model what people prefer.
What does AI red teaming mean?
Red teaming is the practice of having friendly hackers and specialists attack a model deliberately — trying to jailbreak it — so that safety holes surface before someone with hostile intent exploits them.
Will an AI model ever be 100% safe?
A 100% safety guarantee is not something any model can offer. Defence and circumvention keep feeding each other, which makes it an arms race, and that in turn makes ongoing monitoring indispensable.
Why does my AI assistant sometimes balk at harmless questions?
The name for it is "over-alignment": filters set up to block harmful output growing too aggressive, so that benign prompts get declined or answered with excessive caution.
What exactly are AI guardrails?
Guardrails are the filters on either side of a model — one reading what the user sends, one reading what the model returns — that stop malicious material, hate speech or dangerous instructions from getting through to a person. Their role is the last line of defence.
◆

知微

Our work is investigating AI technology and giving ordinary users practical advice on staying safe. Accuracy checks on this guide were last carried out in June 2026. Questions, or something to add? Reach our team or contribute a piece.