一个 AI 模型要做到安全,需要哪些步骤?What Goes Into Making an AI Model Safe?
每一个有能力的模型背后都压着一套严苛的安全机制。本文讲清 2026 年 RLHF、红队测试与护栏如何协同,造出可以放心使用的系统。
Every capable model sits on top of a demanding safety regime. Here is how RLHF, red teaming and guardrails come together in 2026 to produce systems you can rely on.
相对而言,让大语言模型学会东西是比较容易的那一半。难的那一半是让它安全。只要喂给它足够多的网络文本,它在吸收有用内容的同时,也会把谩骂、偏见乃至真正危险的行为模式一并吞下。没有人是特意去教它这些的;它们是默认附带进来的,除非有人主动介入,否则就会一直留在模型里。
这一切出问题时,普通用户会踩到 哪些 AI 风险,我们另有文章讲过。这里要谈的是另一面:公司到底做了哪些具体的事,从源头上阻止这些情况发生。AI 安全——行话叫「对齐」——是这整套工作的统称,它更像层层叠加的多道防线,而不是一剂灵药。
01对齐为什么这么难
模型要先「对齐」——也就是它的目标与人类真正想要的保持一致——才配得上安全这个评价。想象一台动力极强的发动机被装进一辆没有方向盘的车:直线冲刺很惊人,遇到第一个弯道就变成隐患。
难点在于,人类价值并不是一份工整的规格说明书。它们彼此冲突、依赖语境,而且换个提问的人就换个答案。要让机器识别讽刺、权衡两种相互抵触的伦理主张、或者判断界线在哪里,光靠算力不够。它需要长期、有意识的行为工程——而美国国家标准与技术研究院(U.S. National Institute of Standards and Technology)正是想通过其 AI 风险管理框架 把这类问题纳入统一标准。
- 🌐
未经筛选的网络文本
→
🧹
数据清洗
→
🧠
RLHF 对齐
→
🔍
红队测试
→
🛡️
安全发布
02第一阶段:数据的筛选与清洗
在模型开始任何类似「思考」的动作之前很久,安全工作就已经展开,而它的起点是语料。各家实验室从网上抓取数万亿词,但其中几乎没有任何部分是原样进入训练的。
- 有害内容过滤:分类程序自动扫描数据集,剔除仇恨言论、骚扰内容和露骨材料。
- PII 脱敏:脚本识别并遮蔽个人身份信息(PII),例如电话号码、住址和社会安全号码之类。
- 质量启发式规则:低质论坛、垃圾内容以及已知的虚假信息站点会被降权,或者干脆整批剔除。
03第二阶段:RLHF,对齐的引擎
把一个只会猜下一个词的模型,变成愿意努力做到有用、诚实、无害的东西,靠的就是 RLHF,即「基于人类反馈的强化学习」。在整条流水线里,这是外界听说最多的环节,哪怕这四个字母对他们毫无意义。
RLHF 的内部流程
- 监督微调(SFT):由人工撰写数千段示范对话,模型学着模仿这些高质量回复。
- 奖励建模:针对同一条提示词生成多个回答,由人按优劣排序,模型从中学会人类偏好长什么样。
- PPO 优化:训练让模型在那项奖励上尽量拿高分,实际效果就是把人类对安全与实用的偏好内化进去。
04第三阶段:刻意去攻破模型
对齐做完之后,下一步就是刻意搞破坏。实验室组建红队——由善意黑客、领域专家,有时还包括社会学家组成——他们的全部任务,就是在意图更差的人之前先找到弱点。
这也不是随性的活动。NIST 在其关于 对抗性机器学习 的报告中,给出了攻击类别的正式分类法;而由政府支持的 美国 AI 安全研究所 如今会在新模型发布前,与主要实验室共同开展其中一部分测试。
越狱尝试
提示词注入
专家级测试
慢热式诱导
05第四阶段:为运行中的模型架设护栏
模型训练得再充分,一旦真正对外提供服务,仍然需要一张安全网。护栏是对话进行时从外部实时监控的独立系统,与模型在训练中学到的东西彼此分开。
- 1
输入过滤器
在模型看到之前,先检查你输入的内容是否带有恶意意图、PII 或违禁话题。
2
系统提示词
由开发者写在后台的指令,规定助手的角色定位与不可逾越的边界。
3
输出过滤器
检查模型生成的回复,在展示给你之前确认其中没有有害或不良内容。
06第五阶段:上线之后继续盯
上线并不是终点。模型一旦公开,每天都会冒出新的试探方式。因此,实验室团队会研究匿名化的使用模式,趁新式越狱手法或行为漂移尚未扩散时就把它发现。
发现新问题时,往往不需要重训整个模型。团队通常只要调整系统提示词,或者给奖励模型一点推力,让新发现的坏行为代价更高即可。
07今天的安全手段做不到什么
尽管前面讲了这么多,永久的安全并不存在。它的形态是一个循环:防守建起来,有人绕过它,漏洞被补上,循环再从头开始。各国政府大体已经接受,这是一场长期军备竞赛,而不是一项能完工的任务——英国 AI 安全研究所 的存在,基本就是这个判断的产物。
还有一种容易被人忽略的失效模式值得一提:「过度对齐」。为了拦住有害内容,安全过滤器有时会矫枉过正,结果再正常不过的请求被拒绝,回答也变得含糊到毫无用处。如果你想补一补以上内容的基础概念,我们那些 写给初学者的 AI 指南 是很好的下一步。
08读者最常问的问题
AI 安全中最关键的技术是哪一项?
AI 红队测试是什么意思?
AI 模型有可能做到 100% 安全吗?
为什么我的 AI 助手有时连无害的问题都不肯答?
AI 护栏具体指什么?
Getting a large language model to learn is, by comparison, the straightforward half of the job. The hard half is making it safe. Show it enough text from the web and it will soak up abusive, prejudiced and frankly dangerous patterns side by side with all the useful material. Nobody intends to teach it those things; they arrive by default, and they stay unless somebody steps in.
The dangers ordinary people run into once that goes wrong are covered elsewhere on this site. Our subject here is the reverse angle: the concrete steps a lab takes to stop those failures arising at all. AI safety — alignment, in the jargon — is the label for the whole undertaking, and it behaves less like one remedy than like several defences stacked on one another.
01Why Alignment Is So Hard
A model earns the label safe only after it is "aligned" — that is, once its objectives line up with what people actually want. Picture a very strong engine bolted into a car with no steering wheel. Straight-line speed is spectacular; the first corner turns it into a hazard.
The difficulty is that human values do not arrive as a tidy specification. They conflict with one another, they depend on context, and they shift depending on whom you ask. Getting a machine to register sarcasm, to weigh two competing ethical claims, or to sense where a boundary lies takes more than processing power. It calls for deliberate, continuous behavioural engineering — and that is the kind of difficulty the National Institute of Standards and Technology, the U.S. standards body, sets out to corral in its AI Risk Management Framework.
- 🌐
Unfiltered Web Text
→
🧹
Cleaning the Data
→
🧠
Alignment via RLHF
→
🔍
Red Team Testing
→
🛡️
Safe Release
02Phase 1: Sorting and Cleaning the Data
Long before a model engages in anything resembling "thought", the safety effort is already under way — and it begins with the corpus. Labs pull trillions of words from the web, yet hardly any of it enters training untouched.
- Toxicity Filtering: classifier programs sweep the datasets automatically, stripping out hate speech, harassment and sexually explicit material.
- PII Redaction: scripts locate and blank out Personally Identifiable Information — phone numbers, home addresses, social security numbers and the like.
- Quality Heuristics: forums of poor quality, spam, and sites known for misinformation get their weight reduced, or are dropped from the set entirely.
03Phase 2: RLHF, the Engine Behind Alignment
The step that converts a bare next-token guesser into something which makes an effort to be helpful, honest and harmless is RLHF — Reinforcement Learning from Human Feedback. Of all the stages, it is the one outsiders hear about most, even when the four letters mean nothing to them.
Inside the RLHF Process
- Supervised Fine-Tuning (SFT): people write thousands of model conversations by hand, and the system learns to imitate those exemplary replies.
- Reward Modeling: several answers are produced for one prompt, people order them from best to worst, and the system picks up on what those preferences look like.
- PPO Optimization: training then pushes the model to score as high as possible on that reward, which in practice means absorbing human preferences for safety and usefulness.
04Phase 3: Attacking the Model on Purpose
With alignment done, the next stage is deliberate sabotage. Labs assemble red teams — friendly hackers, domain specialists, occasionally sociologists — whose entire task is to locate the weak points ahead of anyone with worse intentions.
Nor is it ad hoc. A formal taxonomy of attack categories appears in the NIST report on adversarial machine learning, and a portion of this testing is now run jointly with the major labs by the U.S. AI Safety Institute, a government-backed body, ahead of any new model release.
Jailbreak Attempts
Prompt Injection Attacks
Expert-Level Testing
Slow-Burn Attacks
05Phase 4: The Guardrails Around a Live Model
A model can be thoroughly trained and still need a safety net once it is serving real traffic. Guardrails are the outside systems that monitor a conversation in real time, operating independently of anything the model absorbed during training.
- 1
Input Filters
Inspects what you typed for hostile intent, PII or prohibited subjects before the model ever receives it.
2
System Prompts
Instructions the developer supplies behind the scenes, setting the assistant's persona and its hard limits.
3
Output Filters
Examines the reply the model produced, checking for toxic or otherwise damaging material before it is displayed to you.
06Phase 5: Watching It After Launch
Launch is not the finish line. The moment a model is public, fresh attempts to probe it appear daily. That is why lab teams study anonymized usage patterns, watching for novel jailbreak methods or slow shifts in behaviour while the problem is still contained.
Fixes for a newly spotted issue often require no full retraining. A team can typically adjust the system prompt, or give the reward model a nudge so that the freshly discovered misbehaviour costs it more.
07What Today's Safety Work Cannot Do
Despite everything described above, permanence is not on offer. The pattern is a loop: a defence goes up, somebody circumvents it, the hole is closed, and the loop begins again. Governments have largely accepted the situation as a standing arms race rather than a task with a finish line — which is essentially why the UK's AI Security Institute exists.
One more failure mode deserves a mention because it is easy to miss: "over-alignment." Safety filters tuned to keep harmful content out can overshoot, at which point ordinary requests get refused and answers turn so hedged that they stop being useful. For the groundwork behind everything discussed here, our AI guides written for beginners make a sensible next read.