AI 的缩放定律到底是什么?So What Exactly Is the Scaling Law in AI?

📈 AI 数学⏱14 分钟阅读📅更新于 2026 年 6 月

AI 变聪明不靠魔法,靠数学。下面这套缩放定律,讲的就是模型规模、训练数据与算力如何组合出智能。

◆知微•📈 AI 数学 · ⏱14 分钟阅读 · 2026 年 6 月 23 日
📈 AI Mathematics⏱ 14 min read📅 Updated June 2026

Nothing magical makes AI smarter — mathematics does. Here are the underlying scaling laws that govern how model size, training data and compute combine to produce intelligence.

◆知微•📈 AI Mathematics · ⏱ 14 min read · June 23, 2026

眼看每年的模型都比上一代明显更强,自然会冒出一个问题——答案并不是工程师把代码写得更漂亮。背后是一条叫「缩放定律」的数学原理。正是它把 AI 从冷门的学术角落推了出来,变成重塑世界的技术。那么这条定律究竟说了什么,又是怎么起作用的?

用大白话说,缩放定律表明:想要更聪明的系统,并不需要什么惊艳的新算法,需要的是更多训练材料、更强算力和更大的模型。换句话说,机器学习遵循「越多越好」的原则——而描述它的曲线既严格又可预测。

01核心思想:AI 最接近「铁律」的东西

想想教孩子识字这件事。一本书只能带来几个词;一千本书能养出一个爱读书的人;一座巨型图书馆再加上一位出色的老师,也许能养出文学天才。AI 遵循的逻辑与此相似——只不过它的数字表现得很听话。

2020 年,OpenAI 的研究者发表了一篇改变格局的论文,题为「Scaling Laws for Neural Language Models」。其结论是:以「预测句子里下一个词的准确程度」来衡量,模型性能会沿一条平滑的幂律曲线变化,而且提前就能推算出来。

真正令人吃惊的是这条曲线是普适的。换架构、往代码里加各种花招都无所谓——把模型规模和性能画成图,同一条可预测的线照样出现。可见性能不是玄学,而是规模层面的工程问题。

02推动 AI 缩放的三根杠杆

想搞懂 AI 为什么变强,就要理解研究者能够调节的三个量。可以把它们看作缩放定律赖以成立的三根支柱:

N
参数(模型有多大)
D
数据集规模(以 token 计)
C
算力(FLOPs)
🧠模型的容量

参数,记作 N

它统计的是网络中人工「神经元」即连接的数量。最早的 AI 只有几百万个,如今的前沿模型有上万亿个。参数量越大,相当于给模型一个更宽敞的「大脑」来存放复杂模式与知识。
📚知识的仓库

数据集规模,记作 D

这指的是系统在训练期间读进去的文本、代码与图像总量。以「token」(大致相当于词块)计数,现代模型的训练语料达到数十万亿级别,来源包括网络、书籍与学术论文。
⚡原始算力

算力,记作 C

这是训练过程消耗的算力总量,以 FLOPs(浮点运算次数)计。要把它交付出来,需要成群的专用 GPU 连续运转数月,耗电量相当于一座小城市。
📈数学表达

幂律本身

按缩放定律,把这三者中的任何一个乘以某个倍数,模型的错误率就会以平滑而规则的方式下降一个可预期的百分比。在计算机科学里,很少有规律能这么稳定地成立。

03这条思路是怎么发展的:从 OpenAI 到 Chinchilla

2020 年 OpenAI 的缩放定律一发表,整个行业就坐不住了。一场竞赛随即开始,每家实验室都想造出最大的模型。数十亿美元砸进超大规模训练,前提假设是「除了规模,别的都不算数」。

到了 2022 年,Google DeepMind 的研究者丢出一篇重磅论文,名为「Chinchilla」。它证明行业此前走错了方向。按最初的定律理解,各实验室把模型做到尽可能大,同时只喂相对很少的数据;DeepMind 则展示了这种做法有多浪费。

Chinchilla 带来了什么改变

Chinchilla 确立的是:数据与模型规模必须同步增长。模型翻倍,训练数据也得翻倍。跳过这一步,得到的会是一个「训练不足」的系统——脑子很大,读的东西却不够把容量用起来。

这一认识重置了优先级。各实验室不再只追规模,转而寻找更多数据,并用更长时间训练更精简、更高效的模型。我们的每周 AI 研究速览追踪着这些高效架构如何改变平衡。

04撞上「数据墙」

整个领域头上压着一个严重难题:数据正在被用尽。缩放定律要求源源不断的文本,而公开互联网上值得读的东西,几乎都已经被写完了。

按专家估算,到 2026 或 2027 年,高质量公开文本的存量就会被耗尽。这就是所谓的「数据墙」。当一条定律预设数据无限、而现实中的数据有限时,曲线会变平吗?AI 的进步会就此停下吗?

数据墙:当人类文本被用尽
  1. 1

    供给枯竭
    公开互联网上每一本高质量图书、每一篇文章、每一个代码仓库和每一条维基页面,都已被 AI 训练吞掉。



    2

    自己造数据
    各实验室开始让能力更强的旧模型出马,批量生成海量高质量「合成」训练材料。



    3

    质量的陷阱
    这里必须谨慎:当 AI 从 AI 的输出中学习,反复迭代会引发「模型坍塌」,答案会变得越来越扭曲、越来越不通。



    4

    注意力转向哪里
    与其硬撞这堵墙,研究者把重心从预训练规模移开,转向「测试时算力」与推理能力。

各公司在用合成数据和私有数据集绕开这堵墙。我们关于最新 AI 研究突破的报道,进一步介绍了针对数据短缺的新架构与合成数据路线。

05测试时算力:另一种缩放

由于训练阶段已经撞上数据墙,业界找到了提升智能的另一条路:测试时算力。

研究者不再只是把最初的训练规模做大,而是在模型回答你的那一刻额外给它算力。想象一场数学考试:普通 AI 想到什么答案就直接写下去;而在测试时扩展过的 AI 会花 30 秒「思考」——检查解题过程、尝试不同的逻辑路径、修正自身错误——然后才给出最终答案。

我们的文章推理型 AI 是什么、如何运作专门拆解了这股「会思考的模型」的趋势。把算力花在推理阶段,让 AI 能解开过去无解的数理、编程和科学问题——这是一条真正全新的、向上而非向外延伸的曲线。

06只靠缩放能抵达 AGI 吗?

压在所有问题之下的那个问题是:这些缩放定律最终会不会通向通用人工智能(AGI)——一种在几乎任何认知任务上都能超过人的系统?

Ilya Sutskever 和 Sam Altman 是「缩放假说」的代表人物。他们的论证是:只要三根支柱(数据、算力、参数)持续增长,AGI 会自行出现——它是数学推演的必然结果,而不是运气。按这种看法,智能不过是规模的函数。

怀疑派则反驳说,收益递减不可避免。他们认为真正的 AGI 需要的是完全不同量级的架构突破——而不只是把更多数据灌进一个更大的网络。许多研究者相信,想弄清楚 AGI 是什么、是否已经实现,唯一务实的路径就是把这些缩放机制吃透,而这场争论至今仍是计算机科学里最激烈的一场。

07读者最常问的问题

AI 缩放定律该怎么定义?
它是一条数学原理:只要三件事增长,神经网络的性能就会可预测地上升——参数数量(模型规模)、训练数据量,以及投入训练的算力。按这个说法,智能在很大程度上由规模决定。
AI 缩放定律是谁做出来的?
OpenAI 的研究者——Jared Kaplan 等人——在 2020 年末发表了奠基性的神经网络缩放定律。他们的工作证明,语言模型的性能沿着一条可提前推算的幂律曲线走,由规模、数据集大小和算力共同决定。
Chinchilla 缩放定律讲了什么?
Chinchilla 缩放定律由 DeepMind 在 2022 年提出,把 OpenAI 的结论又往前推了一步。它指出当时大多数训练都是浪费的——模型相对于其背后的数据而言过大——并主张数据与模型规模必须同步上升,才能把手上的算力用到极致。
AI 的缩放定律是不是已经摸到天花板了?
这个领域已经撞上了「数据墙」:可供训练的高质量人类文本正在变少。研究者的绕行办法是「测试时算力缩放」——让模型在推理阶段「想」得更久——以及制造高质量的合成数据,让曲线继续往上走。
缩放定律对 AI 公司为什么这么有价值?
有了缩放定律,公司在掏数百万美元训练之前,就能知道模型会有多强。用小规模试验先把曲线画出来,就能预测一次耗资数百万美元的大规模训练的结果,省下大量时间和算力。
◆

知微

我们关注的领域是人工智能的数学、技术及其未来形态。本指南的准确性已于 2026 年 6 月复核。对 AI 系统如何搭建有疑问?联系我们的团队,或了解我们到底想做什么。

Watching each year's models arrive markedly sharper than the last raises an obvious question — and the explanation is not simply that engineers write cleaner code. A mathematical principle called the scaling law sits behind it. That principle is what carried AI out of obscure academic corners and turned it into technology that reshapes the world. So what does the scaling law actually say, and how does it operate?

Put plainly, scaling laws show that a smarter system does not require a dazzling new algorithm. What it requires is more training material, more computing power and a larger model. Machine learning, in other words, obeys a "more is better" principle — and the curves that describe it are both strict and forecastable.

01The Central Idea: AI's Closest Thing to an Iron Law

Think about teaching a child to read. One book yields a handful of words; a thousand books produce a devoted reader; a vast library plus an exceptional tutor might produce a literary prodigy. AI follows much the same logic — only with numbers that behave predictably.

Researchers at OpenAI put out a paper in 2020 that changed the field, titled "Scaling Laws for Neural Language Models." Its finding: model performance, gauged by how reliably a system predicts the following word in a sentence, traces a smooth power-law curve that can be forecast in advance.

The startling part is that the curve is universal. Change the architecture, add whatever clever tricks you like to the code — plot model size against performance and the same predictable line reappears. Performance, then, is no mystery; it is a matter of engineering at scale.

02Three Levers That Drive AI Scaling

Grasping how AI improves means understanding the three quantities researchers are able to dial up or down. Call them the supports on which the scaling law rests:

N
Parameters (how big the model is)
D
Dataset size (measured in tokens)
C
Compute (FLOPs)
🧠Capacity of the Model

Parameters, written N

This counts the artificial "neurons," or connections, inside the network. The earliest AI systems had millions of them; frontier models today have trillions. A bigger parameter count amounts to a roomier "brain" for holding intricate patterns and facts.
📚The Knowledge Store

Dataset size, written D

This is how much text, code and imagery the system reads while training. Counted in "tokens" — roughly, word fragments — the training corpora of modern models run to tens of trillions, harvested from the web, books and scholarly papers.
⚡Raw Processing Power

Compute, written C

This is the computational effort the training run consumes, counted in FLOPs (floating-point operations). Delivering it takes enormous clusters of specialised GPUs running for months on end, drawing the electricity of a small city.
📈The Mathematics

The Power Law Itself

According to the scaling law, multiplying any one of these three by some factor brings the model's error rate down by a percentage you can anticipate in a smooth, regular way. Few regularities in computer science hold up as consistently.

03How the Idea Developed: OpenAI, Then Chinchilla

The industry sat up when OpenAI's scaling laws appeared in 2020. A race began, with every lab trying to build the biggest model. Billions went into enormous training runs, on the assumption that nothing but size counted.

Then in 2022 came a bombshell from researchers at Google DeepMind, a paper named "Chinchilla." It showed the industry had been going about things the wrong way. Reading the original laws, labs had built models as large as possible while supplying comparatively little data; DeepMind demonstrated how wasteful that approach was.

What Chinchilla Changed

What Chinchilla established is that data and model size have to grow in step. Double the model and the training data must double too. Skip that and the result is an "under-trained" system: an outsized brain that has not read enough to put its capacity to work.

That insight reset priorities. Rather than chasing size alone, labs hunted for additional data and trained leaner, more efficient models for longer stretches. Our weekly AI research roundup tracks how these efficient architectures are shifting the balance.

04Running Into the Data Wall

A serious difficulty hangs over the whole field: the supply of data is running down. Scaling laws call for ever more text, yet nearly everything on the open internet that is worth reading has already been written.

By 2026 or 2027, on expert estimates, the stock of high-quality public text will be used up. The phrase for this is the "Data Wall." When a law presupposes unlimited data and the supply is finite, does the curve go flat? Does progress in AI simply halt?

The Data Wall: What Happens When Human Text Runs Out
  1. 1

    The Supply Runs Out
    Every high-quality book, article, code repository and Wikipedia page on the open internet has been swallowed by AI training runs.



    2

    Making Data Instead of Finding It
    Labs start putting older, highly capable models to work generating enormous volumes of fresh, high-quality "synthetic" training material.



    3

    The Quality Trap
    Care is essential here: when AI learns from AI output, repeated generations can produce "model collapse," in which answers grow warped and incoherent.



    4

    Where Attention Moves Next
    Rather than trying to break through the wall, researchers are shifting their emphasis away from pre-training scale and toward "test-time compute" along with reasoning ability.

Companies are working around the wall with synthesised data and datasets held privately. Our coverage of the newest breakthrough AI research goes further into the fresh architectures and synthetic-data approaches aimed at the shortage.

05Test-Time Compute: Scaling of a Different Kind

Since the training phase has run up against the data wall, the field has found another route to greater intelligence: test-time compute.

Rather than enlarging the original training run, researchers now hand the model extra compute at the moment it is answering you. Picture a maths exam. An ordinary AI writes down whatever answer surfaces first. One scaled at test time will take 30 seconds to "think" — reviewing its working, trying alternate routes through the logic, fixing its own errors — before committing to an answer.

Our guide to reasoning AI and the way it functions unpacks this move toward models that "think." Extra compute spent at inference lets AI crack mathematical, programming and scientific problems that used to be out of reach — a genuinely new curve that scales upward rather than outward.

06Will Scaling Alone Get Us to AGI?

The question underneath all the others: do scaling laws eventually deliver Artificial General Intelligence (AGI), a system that beats people at essentially any mental task?

Ilya Sutskever and Sam Altman are among the prominent backers of the "Scaling Hypothesis." Their argument runs that AGI arrives on its own if the three supports (Data, Compute, Parameters) keep growing — a mathematical consequence rather than a lucky discovery. On this view, intelligence is simply a function of scale.

Sceptics counter that diminishing returns are inevitable. Genuine AGI, they say, will demand architectural breakthroughs of an entirely different order — not merely more data poured into a bigger network. Plenty of researchers hold that getting to grips with these scaling dynamics is the only practical way to settle what AGI is and whether it has arrived, and the argument remains among the fieriest in computer science.

07Questions Readers Ask Most

How would you define the AI scaling law?
It is a mathematical principle holding that a neural network's performance rises predictably when three things grow: parameter count (model size), training data volume, and the compute devoted to training. Intelligence, on this account, is largely determined by scale.
Whose work produced the AI scaling laws?
Researchers at OpenAI — Jared Kaplan et al. — published the foundational neural scaling laws in late 2020. Their work demonstrated that language model performance tracks a power-law curve you can anticipate, determined by size, dataset volume and compute.
What does the Chinchilla scaling law state?
Introduced by DeepMind in 2022, the Chinchilla scaling law sharpened the OpenAI findings. It showed that most training runs of the day were wasteful — the model was oversized relative to the data behind it — and argued that data and model size must rise together to get the most out of available compute.
Have AI scaling laws started to hit a ceiling?
The field has run into a "Data Wall": high-quality human writing to train on is running short. Researchers are getting around it by "test-time compute scaling" — letting models "think" longer at inference — and by manufacturing high-quality synthetic data to keep the curve climbing.
What makes scaling laws so valuable to AI companies?
With scaling laws, a company can know how capable a model will be before committing millions of dollars to train it. Small pilot experiments let them trace the curve and forecast the results of a giant, multi-million-dollar run, sparing enormous quantities of time and compute.
◆

知微

Our subject is the mathematics, technology and future shape of artificial intelligence. Accuracy for this guide was reviewed in June 2026. Questions about how AI systems are built? Get in touch with our team or read about what we set out to do.