AI 中的 tokenization 是什么?What Is Tokenization in AI?
给 AI 发一条消息,它的第一个动作永远相同:把你的词语切成碎片。这个名为分词(tokenization)的开场动作,藏在 ChatGPT、Claude 和 Gemini 产出的一切内容之下。下文我们会拆解它究竟是什么意思、为何有分量,以及它会如何改变你使用 AI 的方式。
Send a message to an AI and its opening move is always the same: slice your words into fragments. That opening move, tokenization, sits beneath everything ChatGPT, Claude, and Gemini produce. Below we unpack what it really means, why it carries weight, and the difference it makes to how you work with AI.

你或许注意过 AI 的一个怪习惯。让它数一数一个词里有几个字母,答案有时会错;让它把一句话倒过来,它又会结巴。这些并不是笨的表现,而是直接源于一个叫分词的步骤。弄明白 token 是什么,模型的种种古怪之处就一下说得通了。
分词属于那种罕见的概念:它静静躺在你有过的每一次 AI 交流之下,却几乎没人用日常的话把它讲清楚。这个空白今天就来补上。读完之后,你会确切知道自己的文字如何被处理、token 上限为何会塑造你的聊天,以及这些知识如何让你把 AI 用得更好。
01大白话含义
从最底层讲起。计算机无法像人那样阅读文字。hello 这个词到达你大脑时是一个熟悉的整体,计算机遇到的却是一串字符:h、e、l、l、o。在 AI 模型能用语言做任何聪明事之前,它需要一条把人类文字转成可进行数学处理之形式的路径,而这条路径正是分词。
最简单地说,分词就是把文字切成名为 token 之片段的动作。想象一句话被拆成乐高积木,每一块积木都是一个 token。AI 拿起这些积木、逐一检查、弄清它们如何相连,再借助这种理解一块一块地拼出回复。
这与 AI 模型处理语言的整体方式直接相关。若想完整了解 token 生成之后还会发生什么,我们关于AI 如何理解人类语言的指南追踪了从你第一次击键到最终回复的全过程。
02分词为何对你真的重要
你很可能会问:文字在内部怎么切,凭什么要我操心?原因是分词带来三个非常实际、每天都会发生的后果,每一个都触及你使用 AI 工具的日常体验。
1. 上下文窗口:AI 为何会在长聊天里接不上话
每个 AI 模型一次最多容纳一定数量的 token,这个上限名为上下文窗口。把它当作模型的工作记忆。一旦聊天超出这个上限,模型就开始丢弃较早的部分,为新内容腾出空间。
这就是为什么在一段很长的聊天深处,AI 可能像是忘掉了你早先交代过的事。其实什么都没忘,只是那段内容落到了上下文窗口之外。我们关于AI 模型的上下文窗口究竟意味着什么的详细讲解,确切展示了它的运作方式。
2. API 成本:以 token 作为计费单位
当你像开发者搭建应用那样通过 API 使用 AI 时,每个 token 都会被计量——既包括你作为提示词发送的 token,也包括模型作为回复返回的 token。了解 token 数量有助于你起草更精简的提示、可靠地预估费用。为了正好这个用途,本页更靠下的位置放了一个 token 成本计算器。
3. 字母层面任务的准确度
有件事会让大多数人意外:AI 在数单词里的字母方面出奇地弱。问 ChatGPT,strawberry 里有几个 R,回复可能会错。分词是直接原因。模型从来碰不到单个字母,它碰到的是 token。如果 strawberry 被当成一个 token,除非专门为此设计,否则模型没有路径去检查其中的字符。一旦明白这一点,这些“蠢”失误就不再让你惊讶。
03分词实际如何运作
下面是从你向 AI 发出消息那一刻起、逐步发生的过程:
- 1
预处理:清理文字
1 首先是一次轻微清理:处理标点、顾及特殊字符,并把文字规范化,这样无论你恰好怎么打字,处理都保持一致。
- 2
切成 token
2 分词器依据一个学来的词表切割你的文字,通常有 50,000 到 100,000 个可能的 token。the、is、you 这类高频词各占一个 token,较罕见的词和技术术语则被分成子词片段。tokenization 这个词本身可能拆成两个 token:token 和 ization。
- 3
token ID:附上数字
3 每个 token 随后被换成一个唯一整数,即它的 token ID。于是 hello 可能变成 15496,world 可能变成 995。从这里开始,模型根本遇不到词语,到达它那里的只有这些数字。
- 4
嵌入(embedding):让数字带上含义
4 每个 token ID 接着变成一个稠密的数值向量,名为嵌入(embedding)——一长串数字,编码了该 token 的含义以及它与其他 token 的联系。这些向量才是 transformer 模型实际处理的东西。想更深入,参见我们关于神经网络内部发生了什么的文章。
- 5
生成:一次一个 token
5 生成回复时,模型从不会一下子写出整句话。它预测最可能的下一个 token、把它折进上下文,再预测再下一个,如此往复,直到回复完成。这就是为什么你会注意到答案出现时带有一种轻柔的流式效果。
04动手试试:实时 token 计数器
在下面键入任意一句话,看着它立刻被分成 token。这会帮你直观体会 AI 究竟如何“看”你的文字。
05BPE 是什么?驱动当今分词的算法
GPT-4、Claude、Gemini 等大多数主流 AI 模型,都依赖一种名为字节对编码(Byte Pair Encoding,BPE)的分词技术。名字听着吓人,思路却优雅地简单。
BPE 一开始把每个单独字符当成一个 token,所以 hello 开场是五个 token:h、e、l、l、o。随后它扫描训练数据,寻找最常见的 token 对;如果 t 和 h 经常相邻,就合并成一个 token:th。接着合并下一个最常见的对,再下一个,重复成千上万次,直到词表达成目标大小,通常约 50,000 到 100,000 个 token。
结果是这样一套词表:高频英文词各占一个 token,而罕见词、技术行话以及来自较不常见语言的词会被分成更小的子词片段。这解释了下面这些情况:
| 词 / 短语 | 约多少 token | 原因 |
|---|---|---|
| the | 1 | 极其高频,所以赢得自己的专属 token |
| hello | 1 | 一个非常常见的英文词 |
| tokenization | 2-3 | 不那么高频,所以分成子词 |
| unbelievable | 3-4 | 由常见前缀和后缀构成的长词 |
| supercalifragilistic | 6-8 | 罕见词,所以分成许多片段 |
| 😊 | 1-3 | 表情符号,往往是多字节,所以可能被拆分 |
| 100 个英文词 | ~133 | 平均比例:一个词约为 1.33 个 token |
这也澄清了一个反直觉的怪现象:在词前面加空格,或使用不寻常的大小写,都可能改变 token 数量。分词器会对这些细节作出反应,方式常常让 AI 新手意外。
token 成本计算器
当你通过 API 使用 AI 时,每个 token 都标了价。用这个计算器预测你的提示在不同模型下的花费:
06语言不同,token 数量也天差地别
分词最重要却最少被讨论的特征之一,是它随语言变化的剧烈程度。大多数主流 AI 模型的词表严重偏向英语,因为英语占了它们训练数据的大头。这会带来真切的日常后果。
同一个意思用英语表达可能要 10 个 token,而用阿拉伯语、印地语或泰语说同一件事可能需要 20-30 个 token。原因在于这些语言预先建好的 token 更少,迫使分词器更频繁地把词切成更小、字符级的片段。
| 语言 | 相对 token 效率 | 日常影响 |
|---|---|---|
| 英语 | 效率最高(基准) | 成本最低,上下文容量最大 |
| 法语 / 德语 / 西班牙语 | 约多 1.1-1.3 倍 token | 成本小幅上升,差别不大 |
| 俄语 / 希腊语 | 约多 1.5-2 倍 token | 明显更贵,上下文更少 |
| 阿拉伯语 / 印地语 | 约多 2-3 倍 token | 每条消息都显著更贵 |
| 泰语 / 日语 / 中文 | 约多 2-4 倍 token | 成本高得多,上下文更快填满 |
这仍是整个 AI 领域持续改进的前线。较新的模型越来越多地用更丰富的多语言数据训练,帮助它们的分词器在各语言间变得更高效。不过就目前而言,任何为非英语受众搭建 AI 应用的人,都应把 token 效率当作一个严肃的成本与性能因素来权衡。
07关于 AI token 的常见迷思,一次澄清
在日常使用中,token 周围笼罩着不少困惑。我们来解决流传最广的几个误解:
理解分词也能澄清,为何 AI 生成的图像与文本的行为如此不同。图像并不像词那样被分词,它们改用视觉块(patch)或潜空间表示。如果这一面让你好奇,我们关于AI 如何把文字变成图像的文章给出了完整图景。
08考考你自己:token 知识小测验
准备好衡量自己学到了什么吗?三道题,即时反馈。
09常见问题
AI 中的 tokenization 是什么意思?
在 AI 里,一个 token 有多大?
tokenization 为何对用户重要?
BPE 分词是什么?
不同语言的分词方式不同吗?
在 AI 中,分词之后会发生什么?
You may have spotted an odd habit in AI. Ask it to tally the letters within a word and the answer sometimes misses; ask it to flip a sentence backwards and it falters. These aren't marks of dimness; they follow directly from a step named tokenization. Grasp what a token is and the model's oddities suddenly fall into place.
Tokenization belongs among the rare concepts lying silently beneath every AI exchange you have ever had, yet almost nobody explains it in everyday words. That gap closes here. By the end, you will know precisely how your words are handled, why token ceilings shape your chats, and how that knowledge lets you use AI to better effect.
01The Plain-English Meaning
Begin at the very bottom. Computers cannot read text as people do. The word hello reaches your brain as one familiar whole; a computer encounters nothing but a chain of characters, h, e, l, l, o. To perform anything clever with language, an AI model first needs a route out of human text and into a form it can treat mathematically, and that route is precisely tokenization.
At its simplest, tokenization is the act of cutting text into fragments named tokens. Picture a sentence pulled apart into LEGO bricks; each brick is one token. The AI lifts these bricks, inspects them, works out how they connect, and draws on that grasp to assemble a reply, brick by brick.
This ties directly to the broader way AI models treat language. For the complete picture of what follows once tokens exist, our guide on how AI grasps human language traces the whole route from your first keystroke to the finished reply.
02Why Tokenization Genuinely Matters to You
You may well be asking why the internal slicing of your text should concern you. The reason is that tokenization carries three very practical, everyday consequences, each touching your daily experience of AI tools.
1. Context Windows: Why an AI Loses the Thread in Long Chats
Every AI model holds a maximum number of tokens at once, a ceiling named the context window. Treat it as the model's working memory. Once a chat outgrows that ceiling, the model begins discarding earlier portions to free space for fresh material.
That is why, deep into a lengthy chat, an AI can seem to lose something you stated early on. Nothing is truly forgotten; that material simply fell outside the context window. Our detailed explainer on what an AI model's context window really means shows exactly how this works.
2. API Costs: The Token as the Billing Unit
When you reach AI through an API, the way developers build apps, each token is metered, both the tokens you send as your prompt and those the model returns as its reply. Knowing the token counts helps you draft leaner prompts and forecast expense reliably. A token cost calculator sits further down this very page for precisely that purpose.
3. Accuracy on Letter-Level Jobs
Here is a fact that catches most people out: AI is surprisingly weak at counting letters within words. Put to ChatGPT how many R's sit inside strawberry and the reply can miss. Tokenization is the direct cause. The model never meets individual letters; it meets tokens. Treat strawberry as one token and the model has no route to inspect the characters unless expressly built for it. Once that is clear, these silly slips stop surprising you.
03How Tokenization Actually Works
Here is the sequence, step by step, from the instant you send an AI a message:
- 1
Pre-processing: tidying the text
1 First comes a light tidying: punctuation is dealt with, special characters are accounted for, and the text is normalised, so processing stays consistent however you happened to type.
- 2
Cutting into tokens
2 The tokenizer cuts your text against a learned vocabulary, commonly 50,000 to 100,000 possible tokens. Frequent words such as the, is, and you each take a single token, while rarer words and technical terms divide into sub-word pieces. The word tokenization itself may split into two tokens, token and ization.
- 3
Token IDs: attaching numbers
3 Every token is then swapped for a unique integer, its token ID. Thus hello might become 15496 and world might become 995. From here the model never meets words at all; only these numbers ever reach it.
- 4
Embeddings: turning numbers into meaning
4 Each token ID next becomes a dense numerical vector named an embedding, a long list of numbers encoding the token's meaning and its ties to other tokens. These vectors are what the transformer model actually treats. To go deeper, see our piece on what occurs inside a neural network.
- 5
Output: producing tokens in sequence
5 When a reply is generated, the model never writes a full sentence at once. It forecasts the likeliest next token, folds it into the context, forecasts the one after that, and so on until the reply is complete. That is why you notice a gentle streaming effect as the answer appears.
04Try It: A Live Token Counter
Type any sentence underneath and watch it divide into tokens immediately. It builds an intuitive sense of how the AI actually sees your words.
05What Is BPE? The Algorithm Driving Today's Tokenization
Most leading AI models, GPT-4, Claude, and Gemini among them, rely on a tokenization technique named Byte Pair Encoding (BPE). The name sounds daunting, yet the notion is elegantly plain.
BPE begins with each single character as its own token, so hello opens as five tokens, h, e, l, l, o. It then scans the training data for the commonest token pair; should t and h frequently sit together, they merge into one token, th. The next commonest pair then merges, and the next, repeating thousands of times until the vocabulary reaches its target size, usually about 50,000 to 100,000 tokens.
The outcome is a vocabulary in which frequent English words each take a single token, while rare words, technical jargon, and words from less common languages divide into smaller sub-word pieces. This explains the following:
| Word / phrase | Approx. tokens | Reason |
|---|---|---|
| the | 1 | Exceptionally frequent, so it earns its own token |
| hello | 1 | A highly common English word |
| tokenization | 2-3 | Less frequent, so it divides into sub-words |
| unbelievable | 3-4 | A lengthy word built from common prefixes and suffixes |
| supercalifragilistic | 6-8 | A rare word, so it divides into many pieces |
| 😊 | 1-3 | An emoji, often multi-byte, so it may split |
| 100 words of English | ~133 | Average ratio: one word is roughly 1.33 tokens |
This also clarifies a counterintuitive quirk: inserting spaces ahead of words or using unusual capitalisation can shift the token count. The tokenizer reacts to these details in ways that often surprise newcomers to AI.
Token Cost Calculator
When you reach AI through an API, every token carries a price. Use this calculator to forecast what your prompts cost across various models:
06Languages Differ, and Token Counts Differ Wildly
Among tokenization's most important yet least discussed traits is how sharply it varies by language. The vocabularies of most leading AI models lean heavily toward English, since English made up the bulk of their training data. That carries genuine everyday consequences.
The same thought expressed in English may take 10 tokens, while Arabic, Hindi, or Thai may need 20-30 tokens to say the identical thing. The cause is a thinner set of pre-built tokens for those languages, forcing the tokenizer more often to break words into smaller, character-level fragments.
| Language | Relative token efficiency | Everyday impact |
|---|---|---|
| English | Most efficient (the baseline) | Lowest cost and greatest context capacity |
| French / German / Spanish | Roughly 1.1-1.3 times more tokens | A small cost rise, a minor difference |
| Russian / Greek | Roughly 1.5-2 times more tokens | Noticeably costlier, with less context |
| Arabic / Hindi | Roughly 2-3 times more tokens | Markedly costlier with each message |
| Thai / Japanese / Chinese | Roughly 2-4 times more tokens | Far higher cost, and context fills sooner |
This remains an active front for improvement across the AI field. Newer models increasingly train on richer multilingual data, which helps their tokenizers grow more efficient across languages. For now, though, anyone building AI apps for a non-English audience should weigh token efficiency as a serious cost and performance factor.
07Common Myths About AI Tokens, Cleared Up
A good deal of confusion surrounds tokens in everyday use. Let us settle the most widespread misunderstandings:
Grasping tokenization also clarifies why AI-made images behave so differently from text. Images are not tokenized like words; they use visual patches or latent-space representations instead. If that side intrigues you, our article on how AI turns text into images gives the full picture.
08Test Yourself: A Token Knowledge Quiz
Ready to gauge what you have picked up? Three questions with immediate feedback.