YouTube 的 AI 推荐是怎么跑起来的?What Makes YouTube's AI Recommendations Tick?
本来只想看一条 5 分钟的视频。两小时过去,你已经在昨天还从未听说过的领域里连看了十七条。这不是意外——而是一套 AI 系统正在精准执行它被设计出来要做的事。下面就用大白话讲清它的运作方式。
The plan was one five-minute clip. Two hours on, you are seventeen videos into a subject that was unknown to you yesterday. None of that is chance — it is an AI system performing precisely the job it was designed for. Below is how the machinery works, explained in plain English.

你只想看一眼修水龙头的小短片。结果一抬头,晚上没了,人已经掉进苏联潜艇史里,而且完全想不起是怎么滑进去的。很熟悉?那你早就遇上过互联网上最强的推荐引擎——只是从没抓到它工作的样子。
那究竟是怎样的机制在起作用?没有魔法,也没有谁坐在某处替你挑片子。真正存在的是一整套机器学习模型:它们观察你的行为,从几百万有同样行为的人身上学习,并不断重新估算「再来一条视频能留住你」的概率。本文就来撬开这个黑箱,说清一条视频是怎么挤进你的首页、接下来的播放队列或者 Shorts 那一排的。
往下讲之前,有个区分值得先划清楚,因为两个概念总被混为一谈。YouTube 里并不是每个自动决策都算深度学习意义上的「AI」;相当一部分只是基于规则的机械处理,比如把你说都没说过的那种语言的视频挡在外面。AI 到底在哪里结束、自动化从哪里开始,我们这篇 AI 与自动化的分界线在哪 讲得很细。而推荐引擎本身则是货真价实的机器学习,训练它用的是天文量级的观看数据——几十亿次观看会话,多到难以想象。
01说白一点:它学的是你的行为,不是你的说法
把这套系统想象成一个听不见你的偏好、也看不见你的用意的家伙:它只记录行为。每一次点开缩略图都算数。快进前停留的那几秒也算,一打开就退出的短片、留下的评论、敲进搜索框的词,全都算。单独一项只是噪声。可当你在这个平台上累积了几百小时,它们会拼出一幅出人意料地忠于你的喜好画像。
这个套路并不为 YouTube 独有,很多 AI 系统都用同一招。读过我们那篇 AI 怎么从照片里认出人脸 的读者会认出这个形状:把某个杂乱而属于人的东西——在人脸那个例子里就是人脸——转成一组整齐、可比较可测量的数字。观看行为受到的待遇完全一样。观众和视频都被压成一长串数字,这串数字叫「嵌入」(embedding),本质上就是内容与品味的数学指纹。
相似的东西会被有意安排在一个专为此构建的庞大数学空间里彼此靠近。意式咖啡机的视频聚在同一片区域;看过十来条这类视频的人,其嵌入也会落在附近。推荐到这一步就变成一件事:去找离你最近的那些嵌入。如果你还不了解「训练」一个系统这样做事是什么意思,我们这篇 什么是机器学习、它是怎么训练的 会用大白话把基础讲一遍。
02从打开 App 到拿到首页:YouTube 的决策路径
从你点开应用,到一份首页递到眼前,中间走的是这样一条路线:
03动手试试:像算法那样给一条视频打分
每个信号各占多少分量,这事值得亲自拨一拨。按下方的按钮,就能看到不同类型的观看行为如何让一条视频的得分上下摆动,用的是一版大幅简化的替身,模拟 YouTube 给视频排序的方式。
04底层到底是什么:神经网络与 Transformer
排序这一步底下,压着一种常被称作「双塔」模型的深度学习架构。第一座塔根据观看历史,得出观众本人的数值画像;第二座塔根据标题、简介、缩略图、音频以及其他观众的反应,得出视频的数值画像。每当一位真实观众真的喜欢上一条真实视频,训练就会把这两座塔在数学空间里拉得更近。
在较新的版本里,YouTube 的推荐系统也转向了基于 Transformer 的架构。这正是驱动现代聊天机器人和翻译工具的那一族模型,它让系统把你的观看历史当作有序的序列来读,而不是一堆随机堆放的东西。Transformer 在意的是顺序和上下文——就像它会跟踪句子里的词序——而不是把你最近看过的五十条视频塞进一个无序袋子。想了解更技术性的全貌,可以看我们这篇 Transformer 模型究竟是什么,里面拆解了注意力机制的实际运作。
这里和大语言模型有个巧妙的对应。我们那篇 AI 如何定下下一句话 讲的是:模型根据前面出现的一切,预测句子里接下来的那个词。YouTube 的序列模型在结构上做的是同一件事——根据你已经看过的那串视频,预测你接下来会打开哪一条,只是训练它的是观看模式而不是语言。这一切都离不开海量训练数据,我们那篇 为什么 AI 训练需要这么多数据 解释了原因,毕竟 YouTube 的模型每天要从几十亿次观看会话产生的信号中学习。
| 组成部分 | 它的作用 | 日常类比 |
|---|---|---|
| 嵌入层 | 把观众和视频转成可以相互比较的数值向量 | 差不多相当于把品味画成地图上的一个坐标 |
| 双塔模型 | 把观众嵌入与视频嵌入配对 | 差不多相当于媒人拿着两份资料比对合不合拍 |
| 注意力 / Transformer 层 | 判断过去哪些视频对下一次预测最关键 | 差不多相当于回忆刚聊完的一整段话,而不只是记住一个词 |
| 重排序层 | 在原始得分之上叠加多样性、新鲜度和政策规则 | 差不多相当于编辑把关,确保最终清单不会重复单调 |
05推荐系统出现在 YouTube 的哪些角落
说它只是一个功能,实在小看它了。平台摆到你面前的几乎每一个界面之下,都有同一台引擎在运转。
首页信息流
接下来的播放与自动播放
Shorts 信息流
搜索联想
通知提醒
热门与探索
06它到底有多准?又在哪里会失手?
在预测你会点什么、会停留多久这件事上,这套系统确实相当擅长——这正是它被优化的指标——效果也看得见。但「擅长预测点击」并不等于「对你总是有益」,而且这个模型有明显的盲区。
算法会卡住的地方:
- ✗
冷启动问题
✗ 全新的观众和刚上传的视频都没有历史可供模型学习,因此早期推荐只能靠大众流行度做粗略猜测,谈不上贴合你的口味。
- ✗
困在信息茧房里
✗ 因为每一条推荐都在重复你已经看过的东西,信息流会在你毫无察觉的情况下收窄成回音室——除非你主动去找陌生的话题。
- ✗
标题党的引力
✗ 夸张的缩略图和标题能在短期内抬高点击率,即使视频本身让人失望;在观看时长数据把方向拉回来之前,短期信号已经先被带偏了。
- ✗
兴趣已经变了
✗ 当你的兴趣比观看历史变化得更快,模型会在你早已转向别处之后,继续推送那个被放弃的旧爱好。
- ✗
一个账号,几个人共用
✗ 一个账号被全家共用,会产生混乱、常常互相矛盾的信号,因为模型同时在试图讨好好几个不同的人。
07隐私与伦理:尚未有定论的问题
当对人类行为的预测达到这种水准,争论的范围就大大拓宽了——某一份首页准不准,已经不再是全部问题:
控制权并没有完全被拿走。暂停或清除观看历史、把某条内容标记为「不感兴趣」,或者选择「不要推荐这个频道」——每一种操作都会把新的、有意识的信号送回系统,把算法往你更希望的方向推。
08读者问得最多的问题
YouTube 的 AI 推荐具体是怎么运作的?
YouTube 的算法靠哪些信号来推荐视频?
为什么我会收到从没搜过的内容推荐?
推荐可以重置或者改善吗?
YouTube 的推荐算法和聊天机器人的 AI 是同一回事吗?
为什么 Shorts 的推荐方式和长视频不一样?
可以彻底关掉 YouTube 的 AI 推荐吗?
A quick clip on repairing a dripping faucet was all you wanted. Next thing you know, the evening has vanished and you are deep in Soviet submarine history, with no clear memory of how the descent began. Familiar? Then you have already encountered the internet's most powerful recommendation engine — you simply never caught it in the act.
So what is actually going on? No wizardry is involved, and nobody sits in a room choosing clips on your behalf. What exists is a stack of machine learning models: they observe your behaviour, learn from the millions of other people behaving the same way, and keep recalculating the odds that one more video will hold your attention. This piece pries open that black box and traces how a video earns its place in your Home feed, the Up Next queue, or the Shorts row.
One distinction deserves drawing before we continue, since two separate notions get blurred together constantly. Not every automated decision inside YouTube qualifies as "AI" in the deep-learning sense; a fair amount is plain rule-based plumbing, such as hiding content in a language you have never once watched. Our breakdown of where AI ends and automation begins marks that boundary precisely. The recommendation engine, by contrast, is the real thing: machine learning trained on a volume of viewing data — billions of watch sessions — that is hard to picture.
01The Plain Version: Behaviour Teaches It, Not Declarations
Think of the system as deaf to your preferences and blind to your intentions; only conduct gets recorded. Every tap on a thumbnail registers. So does the gap in seconds before you hit skip, the clip you exit at once, the remark you leave behind, the phrase you type into search. One of those on its own is noise. Stack them across hundreds of hours spent on the platform, though, and they resolve into an unexpectedly faithful portrait of what you like.
The trick is not unique to YouTube; plenty of AI systems lean on the same one. Anyone who has read our explainer on how AI picks out faces in photographs will recognise the shape of it: an unruly, very human thing — a face, in that instance — gets turned into a tidy set of numbers suitable for comparison and measurement. Viewing behaviour receives identical treatment here. Viewer and video alike are squeezed into a long sequence of numbers known as an "embedding," effectively a mathematical fingerprint of content and of taste.
Similar items are deliberately placed near one another inside a vast mathematical space built for the purpose. Videos about espresso machines gather in one region, and anyone who has been through a dozen of them acquires an embedding nearby. Recommendation then becomes a matter of hunting for the embeddings closest to your own. Should the idea of "training" a system to behave this way be unfamiliar, the fundamentals are set out plainly in our guide to machine learning and how training works.
02From App Launch to Feed: YouTube's Decision Path
Here is the precise route taken between opening the app and being handed a feed:
03Try It: Score a Video the Way the Algorithm Would
The relative weight of each signal is worth poking at. Push the buttons underneath and watch a video's score swing as different kinds of viewing behaviour get applied, all through a deliberately reduced stand-in for how YouTube orders videos.
04What Sits Underneath: Neural Networks and Transformers
Beneath the ranking stage lies a deep learning architecture commonly called a "two-tower" model. The first tower derives a numerical portrait of the viewer from that person's history. The second derives a numerical portrait of the video from its title, description, thumbnail, audio, and the reactions of other viewers. Training pushes these two towers nearer to each other in mathematical space whenever a genuine viewer genuinely enjoyed a genuine video.
Newer iterations of YouTube's recommendation stack have moved toward transformer-based architectures as well. That is the same family of model sitting behind modern chatbots and translation tools, and it lets the system read your history as an ordered sequence rather than an unordered heap. Order and context are what a transformer pays attention to — much as it tracks word order within a sentence — instead of lumping your last fifty videos into a bag. Our guide to what a transformer model actually is digs into the mechanics of that attention mechanism.
A tidy parallel exists with large language models. As our explainer on how AI settles on its next word describes, such a model forecasts the word that follows, given everything preceding it. YouTube's sequence models do the structural equivalent: forecasting the next video, given the run of videos already behind you — trained on viewing patterns rather than on language. None of it functions without colossal quantities of training data, a point our article on why AI training demands so much data unpacks, since YouTube's models learn from signals produced by billions of watch sessions every day.
| Building Block | Its Job | Everyday Comparison |
|---|---|---|
| The Embedding Layer | Turns viewers and videos into numerical vectors that can be compared | Rather like plotting taste as a coordinate on a map |
| The Two-Tower Model | Pairs viewer embeddings with video embeddings | Rather like a matchmaker weighing two profiles for fit |
| The Attention / Transformer Layer | Decides which earlier videos count most toward the next prediction | Rather like recalling a whole recent conversation instead of a single word |
| The Re-Ranking Layer | Layers diversity, freshness, and policy rules over the raw scores | Rather like an editor checking that the final list avoids repetition |
05Every Corner of YouTube Run by Recommendations
Claiming this is merely one feature undersells it badly. Beneath almost every surface the platform presents, the same engine is at work.
The Home Feed
Up Next and Autoplay
The Shorts Feed
Suggestions in Search
Notification Alerts
Trending and Explore
06Just How Accurate Is It — and Where Does It Slip?
Predicting your clicks and your staying power is something this system does unusually well — that being the metric it was tuned against — and the results are visible. Being skilled at predicting clicks, though, is not the same as being reliably good for you, and there are unmistakable gaps in what the model can see.
Where the Algorithm Comes Unstuck:
- ✗
The Cold-Start Problem
✗ A newcomer and a just-uploaded video bring no history for the model to work with, leaving early suggestions as crude approximations drawn from general popularity instead of your particular taste.
- ✗
Living in a Filter Bubble
✗ Since every suggestion echoes something you have already watched, the feed can narrow into an echo chamber without you noticing — unless you deliberately go looking for unfamiliar subjects.
- ✗
The Pull of Clickbait
✗ Overblown thumbnails and titles can lift click-through rate for a while even when the video underneath fails to deliver, tilting the short-term signals before watch-time figures pull things back.
- ✗
Interests That Have Moved On
✗ When your interests turn over faster than your history does, the model carries on pushing an abandoned hobby long after you have drifted to something else.
- ✗
One Account, Several People
✗ A single account shared across a household generates tangled and frequently contradictory signals, because the model is simultaneously attempting to please several different people.
07Privacy and Ethics: The Unsettled Questions
When prediction of human behaviour reaches this standard, the debate widens considerably — accuracy of any individual feed stops being the whole question:
Control has not been taken away entirely. Pausing or clearing your watch history, marking something as "Not interested," or choosing "Don't recommend this channel" — each of these pushes fresh, intentional signals into the system, steering the algorithm somewhere you would rather it went.