AI 是怎样弄懂人话的?How Does AI Make Sense of What People Say?
你发出一条消息,几秒内一个经过斟酌的回答便返回。可机器到底拿你的话做了什么?在 ChatGPT 或 Claude 的每一次对话背后,都有一套精妙的机制,弄懂它完全不需要计算机专业训练。
You send a message, and within seconds a considered answer comes back. But what is the machine actually doing with your words? Every exchange on ChatGPT or Claude rests on a remarkable mechanism, and following it takes no computer-science training.

回想你上一次往 ChatGPT、Claude 或 Google 的 Gemini 里输入提示:几句话说完,按下回车,回复就像系统真的听懂了你——有时比人说得还贴切。表面之下究竟在发生什么?软件看着一串字母,又如何明白你的意思?
真实的情况既比人们设想的更简单,也更古怪。机器并不以你的方式“阅读”:写下“外面很冷”,它脑海里不会浮现雨天;你提一个问题,它也不会生出好奇。它所做的,是通过接触几乎无法想象的海量文本,精确学会词与词如何关联——而仅凭这一点,就足以支撑一场相当聪明的对话。
01大白话版解释
先从最关键的一点说起:机器并不拥有人类式的语言理解。你看到“狗”这个词,脑中立刻浮现毛茸茸的动物、摇动的尾巴,也许还有儿时宠物的记忆;对系统而言,“狗”唤不起任何画面。它知道的是统计事实:这个词往往紧挨在“吠”“牵引绳”“忠诚”“品种”等词附近。
这听上去像个限制,从某个角度说也确实是,但同一事实又威力巨大。数十亿个训练句子,让模型建出一张极其致密的词语关联地图,足以给出真正显得用心且贴合当下的回复。这不是哲学意义上的理解,却是一个极有说服力的替身。
要真正看清机器如何处理语言,我们得再往下走一层,检视三个承重部件:token、嵌入、transformer。下面的讲解不涉及任何数学。
02什么是 token?(开篇第一步)
文本要能被使用,先得被拆成便于处理的小段,这些小段就是 token,每个大致是一个词或词片段——模型赖以运作的基本单位。
比如“I love learning about AI”可能被切成 ["I", " love", " learning", " about", " AI"]。像“unbelievable”这种更长更绕的词,则可能拆成“un”和“believable”两个 token。切法之所以重要,是因为它决定模型一次能承载多少上下文,这通过它的上下文窗口来衡量。
不妨试试下面这个交互式分词器:
即便一句朴素的话,token 数也出人意料,而这很要紧。模型一次能处理的 token 有上限(即上下文窗口),对话太长时,较早的几轮可能被“遗忘”——不是粗心,而是触及了记忆边界。我们那篇讲模型上下文窗口究竟是什么的完整文章会深入解释这一机制。
03从词到数字(嵌入)
过程在这里变得真正迷人。计算机无法直接使用词,它只处理数字,所以每个 token 都得变成一长串数字,称为嵌入。
精妙之处在于,这些数字绝非随机:它们位于一个巨大的数学空间里,近义的词落在彼此附近。“国王”和“王后”占据相邻的点,“狗”和“猫”挨得很近,“狗”和“望远镜”则相隔遥远。
一个著名例子很能说明问题:取“国王”的嵌入,减去“男人”,加上“女人”,所得的点就非常接近“王后”。模型无人告知,仅凭观察语言,就发现了性别与王室之间的关联。正是这种时刻,让人觉得这项技术近乎魔法。
嵌入也让语境意义成为可能。河边的“bank”与存钱的“bank”不同,由于周围的词会塑造每个嵌入,两者便始终可区分。好奇类似的手法在视觉领域如何运作?我们那篇讲用 AI 从文本生成图像的文章追踪了一个密切相关的概念。
04Transformer:引擎盖下的发动机
接下来是改变整个领域的设计——transformer。它出自 2017 年一篇题为“Attention Is All You Need”的论文,构成 GPT-4、Claude、Gemini 乃至当今几乎所有主流语言模型的核心方案。
更早的系统像慢读者一样逐字读文本,从左到右排成单纵队列,既迟缓,又常在读到结尾时丢开了开头的要点。Transformer 优雅地解决了它:一句话中的每个 token 被同时处理,每个词都一次性与其他所有词相互权衡。
- 1
第一步:分词并嵌入
1 首先句子变成 token,每个 token 再转为数字嵌入,并带上位置信息,于是句首的“狗”与句尾的“狗”得以区分。
- 2
注意力:哪些词更有分量?
2 这是 transformer 的核心。对每个 token,它追问其他哪些词与之最相关。以“The cat rested on the mat because it felt sleepy”为例,代词“it”指向猫,注意力通过同时权衡所有词语关联来敲定答案。
- 3
层级:纵深处理
3 当今模型把 transformer 层堆得很高,有时超过 100 层,每层都让结果更精确。浅层捕捉语法与句法,深层则触及意义、语境,乃至反讽或对比这类微妙之处。
- 4
输出:预测下一个 token
4 最后一步概念上很简单:在处理完一切之后,模型预测最可能紧随其后的 token,再一个 token 接一个 token 重复,直到答复完成。同一个问题问两次,措辞可能因此略有不同。
想了解更完整的内部构造,可以接着读我们那篇讲神经网络内部如何运作的文章。
05注意力:模型挑选焦点的方式
注意力机制堪称现代 AI 最核心的突破,下面用一个例子把它讲具体。
看这句:“奖杯塞不进那只箱子,因为它太大了。”这里的“它”是奖杯还是箱子?读者立刻知道是奖杯,因为正是它的尺寸挡住了自己;但这需要对整句推理,而不只看相邻的词。
这就是注意力的本质:句中每个词都对其他每个词赋予一个权重,模型实际上在追问那个词对理解本词有多大帮助。结合全句,“它”与“奖杯”牢牢绑定,因为这是最合理的指代;那些曾让更简单的旧系统束手无策的歧义,如今就这样被化解。
正是这套机制,也支撑起实用的客服对话。听到“我的订单至今还没送到,我感到很沮丧”,模型聚焦“至今还没送到”(事实性故障)、“沮丧”(情绪)与“我的”(个人情形)。想看它在业务中的实际用法,可见我们那篇讲助力客服的 AI 工具的文章。
06神话对照现实:AI 在语言上的能与不能
机制既已在手,下面来拆解人们对机器语言最常抱有的误解。
这在日常中意味着什么
弄懂机制不只是学术问题,它会改变你使用工具的方式。一旦明白系统做的是高级模式匹配而非真正推理,几个熟悉的怪现象就都讲得通了:
| 情形 | 为何发生 | 该怎么办 |
|---|---|---|
| 满怀把握地给出错误答案 | 它的预测偏向听起来最可能、而非真正正确的回复 | 任何要紧的事实都用自己的来源核实 |
| 模糊问题的重点被错过 | 缺少语境时,匹配会退回最常见的解读 | 把提示写具体,并提供背景 |
| 长对话里较早的几轮被遗忘 | 上下文窗口只容纳固定数量的 token | 换个新话题另开聊天,或把前文压缩概括 |
| 含蓄的反讽没被察觉 | 语气与意图难以仅凭文本 token 编码 | 直说无妨;如果是在开玩笑,就点明 |
| 每次尝试的回复都略有不同 | token 预测在一个名为温度(temperature)的随机性控制下进行 | 并无异常;需要时可要求保持一致或重新生成 |
07自测:一个小测验
简短确认这些概念是否真正落地:三个问题,没有陷阱,只看是否真懂。
08常见问题
AI 凭什么听懂人话?
在语言模型里,token 究竟是什么?
意义是真的被理解了吗?
用大白话说,NLP 是什么?
为什么笑话和反讽有时会跑偏?
在 AI 里,transformer 是什么?
Recall the last prompt you fed into ChatGPT, Claude, or Google's Gemini. A handful of sentences went in, you pressed enter, and the reply came back as though the system truly followed you — on occasion more aptly than a person would. What is really taking place under the surface? How does software look on a run of letters and grasp what you intended?
The truthful account is at once plainer and odder than people assume. A machine does not "read" in your fashion. Put "it's cold outside" before it and no rainy scene appears; ask a question and no curiosity stirs. What it has done, by meeting an almost unimaginable volume of text, is learn precisely how words connect, and that alone turns out to support a strikingly intelligent dialogue.
01The Plain-Language Account
Begin with the single most crucial fact: human-style language comprehension is not what a machine has. When you meet the word "dog," the mind at once summons a furry animal, a moving tail, perhaps a childhood pet. For the system, "dog" evokes no image at all. What it knows is statistical: the word tends to sit near others such as "bark," "leash," "loyal," and "breed."
That reads as a limit, and in one sense it is, yet the same fact carries enormous power. Billions of training sentences give the model so dense a map of word connections that it can issue replies which feel genuinely considered and suited to the moment. This is not philosophical understanding; it is a highly persuasive stand-in.
To see how the machinery really handles speech, we have to descend one floor and examine the three load-bearing parts: tokens, embeddings, transformers. No mathematics is involved in what follows.
02What Is a Token? (The Opening Stage)
Before your text can be used at all, it has to be broken into pieces easy to handle. Those pieces are tokens, each one roughly a word or a word fragment — the elementary currency the models operate in.
Take "I love learning about AI," which may be cut as ["I", " love", " learning", " about", " AI"]. A longer, knottier word such as "unbelievable" can split across two tokens, "un" and "believable." The cut matters because it sets how much surrounding text the model can carry at once, measured through its context window.
Give the interactive tokeniser below a try:
Even a plain sentence yields a surprising count, and that matters. A model can process only so many tokens at a time (its context window), so a very long exchange may cause earlier turns to be "forgotten" — not from carelessness but from reaching the memory boundary. Our full piece on what a model's context window really is explains the mechanism in depth.
03From Words to Numbers (Embeddings)
Here the process grows genuinely striking. Words themselves are unusable to a computer, which deals only in numbers, so each token has to become a long numeric array called an embedding.
The subtle part is that these numbers are anything but random. They sit inside a vast mathematical space in which near-synonyms land near one another. "King" and "Queen" occupy neighboring points; "Dog" and "Cat" lie close; "Dog" and "Telescope" are far removed.
A famous illustration drives it home: take the "King" embedding, remove "Man," add "Woman," and the resulting point sits very near "Queen." The model has uncovered the tie between gender and royalty without being told, merely by watching language. Moments like these are where the technology feels close to magic.
Embeddings are also what allow contextual sense. "Bank" by a river differs from a "bank" that holds savings, and because surrounding words shape each embedding, the two stay distinguishable. Curious how a comparable trick works in a visual setting? Our piece on producing images from text with AI traces a closely related idea.
04The Transformer: Engine Beneath the Bonnet
Now comes the design that altered the field — the transformer, set out in a 2017 paper titled "Attention Is All You Need." It forms the central plan behind GPT-4, Claude, Gemini, and effectively every leading language model now in use.
Earlier systems met text in single file, word following word from left to right as a slow reader does. That was sluggish, and points made early had often slipped away by the closing words. Transformers answered the problem elegantly: every token in a sentence is handled together, so each word is weighed against every other at one time.
- 1
Stage One: Tokenise and Embed
1 First your sentence becomes tokens, and every token turns into a numeric embedding, position included, so "dog" opening a sentence is distinguished from "dog" closing one.
- 2
Attention: Which Words Carry Weight?
2 Here lies the transformer's heart. For each token it asks which others bear most on it. Take "The cat rested on the mat because it felt sleepy": the pronoun "it" points to the cat, and attention settles the question by weighing every word tie together.
- 3
Layers: Processing in Depth
3 Present-day models pile transformer layers high, at times beyond 100, each one sharpening the result. Early layers catch grammar and syntax; deeper ones reach meaning, context, and even shades of irony or contrast.
- 4
Output: Forecasting the Next Token
4 The last move is conceptually simple: given all it has processed, the model forecasts the token likeliest to follow, then repeats the move token on token until the reply is done. Ask twice and the wording may shift slightly for exactly this reason.
Readers wanting the fuller internal anatomy can continue with our explainer on the workings within a neural network.
05Attention: The Means by Which a Model Picks Its Focus
The attention mechanism has a fair claim to being modern AI's central breakthrough, so an example will sharpen it.
Consider this: "The trophy would not fit in that suitcase because it was too large." Is "it" the trophy or the case? A reader knows at once — the trophy, since its size is what blocked it — yet that calls for reasoning across the whole line rather than the neighboring word.
That is attention in essence: every word in a line receives a weight toward every other, the model in effect asking how far that word helps interpret this one. "It" binds strongly to "trophy" as the most sensible referent across the full line, which is how ambiguity that defeated older, plainer systems now gets resolved.
The very mechanism also powers useful customer-service dialogue. Hearing "my order still has not shown up, and I am feeling frustrated," the model centers "still has not shown up" (the factual fault), "frustrated" (the feeling), and "my" (a personal case). Watch the business uses in our piece on AI tools that aid customer service.
06Myth Against Reality: The Bounds of AI With Language
With the mechanics in hand, let us dismantle the misunderstandings people most often bring to machine language.
What This Means Day to Day
Grasping the mechanism is no mere academic matter; it reshapes how you use the tools. Once the system is understood as advanced pattern matching rather than true reasoning, several familiar quirks fall into place:
| Situation | Why It Occurs | What to Do |
|---|---|---|
| A wrong answer delivered with assurance | The forecast favors the reply that sounds likely, not the reply that is correct | Check any consequential fact through your own sources |
| The point of a vague question is missed | Without context, matching falls back on the commonest reading | Make the prompt concrete and supply surroundings |
| Early turns in a long exchange are forgotten | The context window admits only a fixed number of tokens | Open a fresh chat for a new subject, or compress earlier context |
| Subtle sarcasm passes unnoticed | Tone and intent resist being encoded in text tokens alone | Speak plainly; if a remark is in jest, label it |
| Replies differ slightly on each attempt | Token forecasting runs under a randomness control named temperature | Nothing is amiss; request consistency or a regeneration when needed |
07Check Yourself: A Brief Quiz
A short confirmation that the ideas have landed: three questions, no traps, only genuine grasp.