AI 究竟怎样从人类反馈中学到东西?By What Means Does AI Pick Things Up From Human Feedback?
看一个混乱的文本预测器,如何蜕变为既有用又无害的助手:本文拆解 RLHF、DPO 以及那批不为人知的人工力量,在 2026 年究竟怎样调教机器。
Watch a disordered text predictor turn into an assistant that is both helpful and harmless: here is how RLHF, DPO and the unseen human labour force actually school machines in 2026.
让如今的助手写一首诗、帮忙理顺一段缠成乱麻的代码,或者讲讲量子物理,得到的回答听起来惊人地像人——客气、条理清楚、确实管用。可藏在它底下的东西,起初完全不是这副模样。未经打磨的网络曾经是个狂放、不靠谱的文本生成器;放任不管,它同样可能吐出有毒的胡话、带偏见的咆哮,或一堆词不达意的拼凑。是什么把这样一个粗粝的算法,带过鸿沟,变成一个助人而不伤人的助手?完成这场跨越的正是「对齐」,而真正值得追问的是跨越如何发生——人类反馈究竟怎样教会 AI。
到 2026 年,这已不再是计算机科学家才需要关心的问题。随着模型越来越深地进入日常,从客服坐席到判读医学影像,搞懂那套让它们贴合人类价值观的训练,对谁都有好处。本指南会拆解 RLHF 底层如何运转,梳理更新的 DPO 路线,并介绍那批承担教学、却藏在幕后的人力。
01原始模型的麻烦:为什么光靠预训练不够
在能从反馈中学习之前,语言的基本功得先打牢。这项地基在「预训练」阶段完成:模型扫读海量互联网内容——书籍、文章、代码、论坛——只抱着一个窄窄的目标,即「下一 token 预测」,于是它学到,在「天空是」这几个字后面,最可能接上的词是「蓝色」。
一个严重的问题随之而来。互联网充斥着争吵、谣言、辱骂和胡言乱语,而只追统计概率的原始模型,对「真实」「有用」或「安全」没有任何天生的感觉。问它怎么烤蛋糕,也许会得到一份完美食谱,也许是一段蛋糕成精的虚构故事,更糟的是从网络阴暗角落抄来的有害内容。普通用户或许能在推理阶段靠提示工程引导输出,但有用与安全,早在公众接触模型之前就必须烙进它的基础权重里。
研究者给出的答案是一条多阶段流水线,把混乱的文本预测器塑造成有用的助手,而其中最有名、采用最广的一环,便是 RLHF。
02RLHF:衡量一切方法的标杆
基于人类反馈的强化学习(RLHF)是让 ChatGPT 这类模型成为可能的关键突破,它把冷冰冰的统计预测与人类判断的细腻之处接到了一起。过程虽繁复,却可以干净地拆成三个先后相接的阶段。
第一阶段:监督微调(SFT)
先用一份规模较小、经过精挑细选的提示与人工示范答案,对原始模型做微调。外包人员写出一个个「好」对话应有的样子,让模型掌握对话的基本形态(用户:X;助手:Y),并具备一条能力底线。即便如此,它大多仍在模仿写示范的人,底层的偏好并没有真正吃透。
第二阶段:训练奖励模型
真正的「人类反馈」在这一步登场。把一条提示交给 SFT 模型,让它生成一串不同答案——比如同一问题的四种写法——再由标注员逐一阅读,按有用、诚实、无害等标准,从最好到最差排出次序。
成千上万份这样的排序,被用来训练一个独立模型——「奖励模型」。它只有一项任务:猜测人会给眼前的答案打多少分,学会给好回答一个高「奖励分」,给差回答一个低分。
第三阶段:强化学习(PPO)
最初的 SFT 模型再次被放开去生成,只不过这次打分的不再是人,而是奖励模型。用大白话讲强化学习能让最后一步更好懂:就像狗学动作,只要做了合奖励模型心意的事,就能得到一块「零食」(一个高分),随后不断调整内部,好换到更多零食。在近端策略优化(PPO)之下,AI 慢慢学会持续交出高分回答,把人类偏好内化成自己的偏好。
03更新的挑战者:直接偏好优化(DPO)
RLHF 改变了整个领域,但它有个实打实的弱点:想让训练保持稳定极其困难。单独训练奖励模型平添许多层复杂度,PPO 阶段又以反复无常出名;一点小毛病就可能让 AI「钻空子」,刷出怪异重复的文字——数学上分数漂亮,人读着却像垃圾。
直接偏好优化(DPO)于 2023 年提出,到 2026 年已是许多前沿模型的选择,它把这团乱麻一刀劈开。一个数学上的洞见让整套流程大为瘦身:不必另训奖励模型,也不必跑复杂的强化学习,DPO 用一个巧妙的数学手法,直接拿人类偏好数据更新 AI 的策略。
| 比较维度 | RLHF,传统路线 | DPO,现代路线 |
|---|---|---|
| 独立奖励模型 | 需要 | 不需要 |
| 强化学习环节 | 需要,经由 PPO | 不需要 |
| 训练的稳定程度 | 低——容易被钻空子 | 高——普通监督损失 |
| 算力需求 | 极高 | 中等 |
| 所依赖的人工数据 | 对答案的排序 | 对答案的排序 |
去掉奖励模型后,DPO 训练更快、稳得多,也更难被奖励攻击。它说明一件事:教 AI 懂人类口味,未必非要一套重量级强化循环,找对数学框架就够了。
04机器背后的人:标注员是谁
人们谈起 AI 总像在谈魔法,可 RLHF 和 DPO 同样完全建立在体量巨大的人工劳动之上。反馈来自数据标注员、外包人员,并且越来越多来自某个领域的专家。
通用聊天机器人的评分依据详尽得惊人的细则,涵盖事实正确性、语气、安全,甚至细微的偏见;但模型一旦变强,泛泛的标注员就不够用了。一个不会写代码的人,没法可靠地给几份 Python 输出排序,这与科学家如何衡量 AI 有多聪明、靠严谨红队与对抗性测试来检验,是紧密相连的问题。企业如今请来专家级标注员——拥有博士学位的化学家、资深工程师、法律专业人士——在高水准上评判专业模型。
05衡量成效:怎么知道它真学进去了?
RLHF 或 DPO 训练结束后,开发者没法直接问 AI 是否变得有用又安全——它永远会回答「是」。于是他们依靠自动基准与人工评估的组合。
模型会被放进 MMLU 基准等全面测试,确认对齐没有把核心智力磨钝——也就是所谓的「对齐税」;同时还要过「红队」一关,由道德黑客和专家想尽办法诱使它吐出有毒内容或不慎泄密。拒绝得当,或把圈套化解得体,才算通过。
这种严谨之所以要紧,是因为模型如今运行在高风险场景里。看看 AI 在科研中的角色,我们依赖的是高度对齐、忠于事实、不产生幻觉的系统;在实验室里,一个没对齐的模型可能建议危险的化学品配比,或把关键数据读错。
- 🌐用原始互联网数据预训练
→
📝在监督下微调
→
👥由人构建的奖励模型
→
🧠用强化学习调校策略(PPO)
06下一步:AI 给 AI 打分,以及宪法式 AI
人工劳动是对齐中最紧的瓶颈:雇几千名专家给几百万份答案排序,又慢又贵得惊人,而当模型冲向万亿参数,RLHF 流水线也开始吃紧。讽刺的是,前路要靠 AI 自己,去解决 AI 造成的难题。
「宪法式 AI」与「AI 辅助反馈」正被开拓:给一个能力很强的模型一组核心原则(一部「宪法」),让它去批评、改写较弱且未对齐模型的产物,充当人工标注员,凭借自己对安全与有用的理解,大规模提供反馈。让 AI 定义自己的价值观,确实引出真实的哲学疑问;但对于眼下正在打造的超大型下一代模型,这是唯一可行的路。
07常见问题
AI 主要通过什么过程从人类反馈中学习?
RLHF 与 DPO 有什么不同?
AI 为什么非要有人类反馈?
训练 AI 的反馈由谁提供?
AI 能不靠人类反馈就自我对齐吗?
人类反馈会让 AI 变笨吗?
Ask one of today's assistants for a poem, help untangling tangled code, or a tour of quantum physics, and the answer comes back sounding strikingly human — courteous, neatly arranged, useful. What sits underneath, though, began life as anything but. The raw network was once a wild, unreliable generator of text; freed of guidance it might just as readily pour out toxic nonsense, biased tirades or word salad. What carries such a brute algorithm across the gap to an assistant that helps without doing harm? That crossing is the work of alignment, and the question worth asking is how the crossing is made — how, precisely, human feedback teaches an AI.
By 2026 this is no longer a question reserved for computer scientists. As the models work their way deeper into daily life, from support desks to reading medical scans, it pays to understand the training that lines them up with human values. This guide walks through what goes on under the hood of RLHF, traces the more recent DPO approach, and introduces the hidden human labour doing the teaching.
01The Trouble with Raw Models: Why Pre-Training Alone Falls Short
Lessons from feedback cannot begin until the basics of language are in place. That groundwork is laid in "pre-training," as the model sweeps through enormous stretches of the internet — books, articles, code, forums — with one narrow goal, "next-token prediction": it picks up that after the words "The sky is," the likeliest follow-up is "blue."
A serious problem follows. The internet is thick with arguments, falsehoods, abuse and gibberish, and a raw model chasing statistical odds carries no built-in feel for "truth," "helpfulness" or "safety." Ask it how to bake a cake and a flawless recipe may come back — or a made-up tale about a cake that came alive, or, worse still, something harmful copied from a dark corner of the web. End users may try to steer outputs at inference time with prompt engineering, but helpfulness and safety have to be baked into the foundational weights long before the public ever meets the model.
The answer researchers found is a multi-stage pipeline that takes a disordered text predictor and shapes it into a useful assistant, and the best-known, most widely used stage of that pipeline is RLHF.
02RLHF: The Benchmark Every Other Method Is Measured Against
Reinforcement Learning from Human Feedback (RLHF) is the breakthrough that put models such as ChatGPT within reach, joining cold statistical prediction to the subtleties of human judgement. Intricate as it is, the procedure unfolds as three clean, ordered stages.
Stage One: Supervised Fine-Tuning (SFT)
A smaller, carefully chosen set of prompts and model answers written by people is used first to fine-tune the raw model. Contractors draft examples of what a "good" exchange looks like, which gives the model the basic shape of a conversation (User: X, Assistant: Y) and a floor of competence. Even so, it is mostly copying the people who wrote those examples; the preferences underneath have not genuinely sunk in.
Stage Two: Building the Reward Model
This is the point where genuine "human feedback" enters. The SFT model is handed a prompt and told to produce a spread of answers — four distinct takes on the same question, say — and annotators read through them and arrange them from strongest to weakest against measures such as helpfulness, honesty and harmlessness.
Many thousands of such orderings go into training a separate model, the "Reward Model," with one task: guess how a human would score any answer it sees, learning to stamp good replies with a high "reward score" and poor ones with a low mark.
Stage Three: Reinforcement Learning (PPO)
The original SFT model is set loose to generate again, only now the Reward Model, rather than a person, does the marking. Reinforcement learning in plain language makes the last step easier to picture: think of a dog picking up a trick and earning a "treat" — a strong reward score — whenever it pleases the Reward Model, then re-tuning its internals to collect more treats. Under Proximal Policy Optimization (PPO), the AI slowly learns to answer in ways that keep scoring well, absorbing human preferences as its own.
03The Newer Contender: Direct Preference Optimization (DPO)
RLHF changed the field, but it carries a real weakness: holding training steady is enormously hard. A separately trained Reward Model adds layers of complication, and the PPO stage is famously temperamental; minor faults can let the AI "hack" the model, churning out strange, repetitive prose that scores brilliantly in the math and reads like rubbish to people.
Direct Preference Optimization (DPO), released in 2023 and by 2026 the choice across many leading-edge models, cuts through the tangle. A mathematical insight collapses the whole procedure: rather than train a separate Reward Model and run elaborate reinforcement learning, DPO reaches for a neat trick in the math to refresh the AI's policy straight from human preference data.
| Aspect | RLHF, the Traditional Route | DPO, the Modern Route |
|---|---|---|
| A Separate Reward Model | Yes | No |
| Reinforcement Learning Stage | Yes, via PPO | No |
| How Steady Training Stays | Low — open to hacking | High — ordinary supervised loss |
| Compute Demands | Extremely High | Medium |
| Human Data It Draws On | Orderings of answers | Orderings of answers |
Dropping the Reward Model makes DPO quicker to train, far steadier and much harder to reward-hack. The lesson is that a heavyweight reinforcement loop is not strictly required to teach human tastes; the right mathematical framing will do.
04The People Behind the Machine: Who the Annotators Are
AI is often spoken of as though it were magic, yet RLHF and DPO alike rest entirely on very large bodies of human labour. The feedback comes from data annotators and contractors, and, more and more, from specialists in a given field.
General chatbots are scored against extraordinarily detailed rubrics covering factual correctness, tone, safety and even faint biases, but annotators drawn from the general pool no longer suffice once the models grow capable. Outputs of Python code cannot be ranked reliably by someone who cannot code, a problem closely tied to the way scientists measure how intelligent an AI is through disciplined red-teaming and adversarial tests. Companies now bring in expert annotators — chemists with PhDs, senior engineers, legal professionals — to judge specialist models at a high level.
05Judging Results: How Do We Know the Lessons Stuck?
Once RLHF or DPO training is complete, developers cannot just ask the AI whether it has grown helpful and safe; the answer will always be yes. Instead they lean on a mixture of automatic benchmarks and evaluation by people.
The model is put through sweeping tests such as the MMLU benchmark, to confirm that alignment did not blunt its core intelligence — the so-called "alignment tax" — and through "red teaming," where ethical hackers and specialists do their best to bait it into toxic output or careless disclosure. A well-judged refusal, or a graceful handling of the bait, counts as a pass.
Rigor of that kind matters because these models now run in high-stakes settings. Looking at AI's role in scientific research, we depend on systems that are strongly aligned, factual and free of hallucinations; in a laboratory, a misaligned model might suggest a dangerous mixture of chemicals or badly misread crucial data.
- 🌐Raw Internet Pre-Training
→
📝Fine-Tuning Under Supervision
→
👥Reward Model Built by People
→
🧠Policy Tuned with Reinforcement Learning (PPO)
06What Comes Next: AI Grading AI, and Constitutional AI
Human labour is the tightest bottleneck in alignment, since engaging thousands of specialists to rank millions of answers is slow and extraordinarily costly, and the RLHF pipeline strains as models climb toward trillions of parameters. Ironically, the road forward runs through AI itself, set to work on the problems AI created.
"Constitutional AI" and "AI-assisted feedback" are now being pioneered: a highly capable model receives a short set of core principles — a "constitution" — and is told to criticise and rewrite what a weaker, unaligned model produces, standing in for the human annotator and supplying feedback at scale from its own grasp of safety and helpfulness. Letting an AI define its own values raises genuine philosophical questions, yet for the enormous next-generation models now being built, it stands as the only workable route.