AI 是如何决定接下来该说什么的?How Does AI Choose What to Say Next?
看上去像魔法,底层其实是算术。从把你的句子切成 token,到在成千上万个候选里抽出唯一一个,这条完整链路让大语言模型一步步写出像人说的话。
It looks like magic; underneath it is arithmetic. Here is the full chain — from cutting your sentence into tokens to sampling one candidate out of thousands — that lets a large language model produce human-sounding text one step at a time.

向 ChatGPT 或 Claude 抛出一个问题,答案在一两秒内就开始往外冒:通顺、切题、像是想过的。人很容易从中读出意图,仿佛屏幕后面真有个东西在权衡怎么回答。实际上并没有。它既没有在调取记忆,也没有在形成观点——你看到的是一场极其精密的猜词游戏。
把外壳剥掉,问题只剩一句:到目前为止出现的每个词,下一个最可能是哪个? 模型每秒要把这句话问上几十遍,每问一次,句子就长几个字符。下面我们逐层拆开这套机制:token、概率排序、让长文保持连贯的「注意力」技巧,以及决定输出干巴巴还是活蹦乱跳的那个参数。
01核心直觉:把输入法联想放大一万倍
最好用的参照物,是手机输入法上方那条联想栏。你打出「明天见」,它跳出「吧」或「了」——不是因为它懂你的安排,而是这个搭配在你的历史消息里出现过太多次。
大语言模型就是同一套机制,只是音量被拧到了荒谬的档位。把「你个人的聊天记录」换成公开互联网的一大片——书籍、新闻、论坛帖、源代码——这张模式表就大到能覆盖英语以及许多语言的几乎任何说法。所以当你提问时,并不存在「查库」这一步:回答是逐个词拼出来的,每一步都在赌「按它读过的所有东西来看,接下来最顺的是什么」。
02第一步:为什么词必须先变成数
机器完全不知道「苹果」或「跑」指什么,它只认数。所以在任何预测发生之前,你的文字必须先被重新编码成能做算术的东西,这一步就叫分词。
Token 跟词并不一一对应。它可能是一个音节、一个常见前缀,甚至单个字符。比如 unbelievably 常会被切成三块:un、believ、ably。每个 token 都有自己的编号,模型再把这些编号放进庞大的矩阵里做乘法,算出彼此的关系。想看这一步的完整展开,AI 里的 tokenization 是什么 讲了文字如何变成算式。
03第二步:给所有候选词排队
输入处理完之后,模型会去查它在训练中攒下的那些参数,然后列出一张很长的清单:所有可能接下去的候选词,每个都带一个可能性分数。这张清单就是概率分布。
这张图要这样读:如果前面几个词是「She opened the」,模型给「door(门)」的分远高于其他。但它并不会每回都直接取第一名——那样输出会不断打转、乏味至极。采样是从领先的一小撮候选里抽取一个,变化和一些个性正是从这里来的。
04第三步:上下文如何左右选择
上下文是早期模型的软肋。给它一句以「The bank was steep」开头的话,它可能把 bank 当成银行。转折点来自一种叫 Transformer 的架构,其内部的「自注意力」机制让模型在确定含义前,能同时权衡句中所有其他词。
河岸和银行就是靠这一点分开的。steep、money 这类词把概率往不同方向拽,模型顺着这股拉力走。能跨越长距离记住上下文,输出才显得连贯;AI 里的 Transformer 模型是什么 有更细的架构讲解。
05温度值:在稳妥与意外之间的一根旋钮
同一个模型,有时一本正经、干巴巴,有时又俏皮发散——通常就是温度这个参数在起作用。把它想成一根旋钮,控制模型选词时被允许冒多大风险。
- 低温(0.2):模型偏保守,几乎总取概率最高的那个 token。写代码、做事实摘要时更稳,因为把话说对比说得有趣重要。
- 高温(0.8 以上):开始冒险,偶尔会挑一个排名靠后但「气质合适」的词。文字更鲜活、更像人写的——同时胡说八道的概率也跟着涨。
06模型为什么会张口就编
所谓「幻觉」——理直气壮地讲一件彻底错误的事——是对 AI 最常见的不满,而它的根源正好落在选词这一步:模型优化的目标是「像真的」,不是「是真的」。
只要一句假话恰好长在一副常见的句式骨架里,它就能蒙混过关——模型唯一的检验标准就是「听起来对不对」。整个流程里没有任何核实模块,只有一台概率计算器。所以凡是要紧的信息,都值得另找来源核一遍。这也正是 AI 与传统自动化的分界:一个听规则,一个听模式,AI 与自动化有什么区别 有完整说明。
07能力从哪来:训练数据
模型的本事上限,就是它读过的材料的质量上限。碰到完全没见过的领域——比如一本小众工程手册——它就无从判断那个语境下该接什么词。这也是各家愿意砸巨资收集「广而优」语料的原因:吸收的好样本越多,模型在小众话题上的概率估计就越准。AI 为什么需要那么多数据来训练 有更细的展开。
08落到实处的例子:翻译
翻译最能体现这套机制的威力。老一代工具逐个词替换,出来的句子生硬拗口。现在的模型会先读完整句、吃透上下文,再在目标语言里生成「最可能表达同一意思」的说法。
它不是在孤立地处理一个「hello」。语气、正式程度、文化语境都会参与进来,下一个词随之改变:正式场合用庄重的招呼,朋友之间用随口的一句。这种对语境的敏感度,就是现在和上一代翻译工具的分水岭,AI 翻译是怎么工作的 逐步演示了这个过程。
09常见问题
AI 在吐出下一个词时,内部到底发生了什么?
它那些回答背后,有真正的理解吗?
模型为什么会那么自信地讲错话?
token 究竟指什么?
文本吐得有多快?
Ask ChatGPT or Claude a question and an answer starts appearing within a second or two: fluent, on-topic, seemingly considered. It is tempting to read intent into it, as though something behind the screen were weighing up a reply. Nothing of the sort is happening. No memory is being retrieved and no opinion is being formed; what you are watching is an extremely elaborate guessing game.
Strip it down and the whole question is this: given every word written so far, which word has the best odds of coming next? The model asks that over and over, dozens of times a second, and each answer lengthens the sentence by a few characters. Below we take the machinery apart — tokens, probability rankings, the "attention" trick that keeps long passages coherent, and the setting that decides whether the output comes out dry or playful.
01The Core Idea: Autocomplete, Scaled Up
The handiest comparison is the suggestion strip above your phone keyboard. Type "I'll see you" and it offers "tomorrow" or "later" — not because it grasps your plans, but because that pairing has turned up millions of times in messages you wrote before.
A large language model is that same mechanism turned up to an absurd volume. Swap your personal message history for a large slice of the public web — books, news, forum threads, source code — and the pattern table grows vast enough to cover almost any phrasing in English and plenty of other languages. So there is no lookup happening when you ask something. The reply gets assembled word by word, each step betting on the most sensible continuation given everything the model has previously read.
02Phase 1: Why Words Have to Become Numbers
A machine has no idea what "apple" or "run" refers to. Numbers are all it can work with. So before any prediction can happen, your text has to be re-encoded into something arithmetic-friendly — a step called tokenization.
Tokens rarely line up neatly with words. A token might be a syllable, a common prefix, or a lone character. "Unbelievable", for instance, often splits into three: "un", "believe", "able". Every token carries its own ID number, and the model multiplies those numbers through elaborate matrices to work out how they relate. For the deeper version of this step, how text gets turned into arithmetic walks through how text turns into arithmetic.
03Phase 2: Ranking Every Candidate
With your prompt processed, the model consults the parameters it accumulated during training and produces a long list of every word that could plausibly follow, each tagged with a likelihood score. That list is the probability distribution.
Read the chart like this: if the preceding words were "She opened the", the model scores "door" far above everything else. Yet it does not simply take the top entry every time — do that and the text loops and bores. Sampling draws from the small set of leading candidates instead, and that is where variety and a little personality enter the output.
04Phase 3: How Context Steers the Choice
Context was the weak point of early models. Feed one a sentence beginning "The bank was steep" and it might read "bank" as a financial institution. The fix arrived with an architecture called the Transformer, whose "self-attention" mechanism lets a model weigh every other word in the sentence at the same time before settling on a meaning.
That is what separates a riverbank from a savings bank. Words such as "steep" or "money" pull the probabilities in different directions, and the model follows the pull. Holding context across long stretches is what makes the output feel coherent; the architecture explained in depth covers the architecture in more detail.
05Temperature: The Dial Between Safe and Surprising
You may have noticed the same model sounding flat and factual one moment, then loose and inventive the next. A setting called temperature is usually what changed. Think of it as a dial governing how much risk the model may take when picking words.
- Low temperature (0.2): the model stays cautious, nearly always taking the highest-probability token. Good for code and factual summaries, where being right matters more than being interesting.
- High temperature (0.8 and up): it starts gambling, occasionally choosing a lower-ranked word because it fits the mood. The prose gets livelier and more human — and the odds of outright nonsense climb with it.
06Why Models Make Things Up
"Hallucination" — the confident delivery of something flatly untrue — is the most common complaint about AI, and it traces straight back to how the next word gets chosen. What the model optimises for is plausibility, not accuracy.
A fabricated fact can slide through whenever it happens to sit in a familiar sentence shape; sounding right is the only test applied. There is no verification module anywhere in the loop, just a probability calculator. Which is why anything consequential deserves a check against a source. It is also the line between AI and conventional automation: rules drive one, patterns drive the other — where automation ends and AI begins explains the split.
07Where the Skill Comes From: Training Data
A model can only be as capable as the material it learned from. Show it a domain it has never encountered — a niche engineering manual, say — and it has no basis for guessing the next word in that register. Hence the enormous spending on assembling broad, high-quality corpora: the more good examples a model absorbs, the sharper its probability estimates get on narrow subjects. why training sets have to be enormous goes into the detail.
08Seen in Practice: Translation
Translation is where this mechanism looks most impressive. Older tools worked word by word and produced stilted, wooden output. A modern model instead reads the whole source sentence for context, then generates the target-language wording with the highest odds of carrying the same meaning.
It is not deciding how to render a bare "hello" in isolation. Tone, formality and cultural setting all feed in, and the next word shifts accordingly — a stiff greeting in a formal setting, something offhand among friends. That sensitivity to context is what separates current translation from the earlier generation; the pipeline behind machine translation shows the process step by step.