研究者究竟怎样衡量 AI 的智能?How Exactly Do Researchers Measure AI Intelligence?

🧪 AI 评测⏱14 分钟阅读📅更新于 2026 年 6 月

模型据说能通过律师资格考试,还能写出毫无瑕疵的代码——可我们凭什么说它「聪明」?走进实验室看看:2026 年研究者用来给机器智能打分的基准测试、图灵测试和对抗性评估。

◆知微•🧪 AI 评测 · ⏱14 分钟阅读 · 2026 年 6 月 23 日
🧪 AI Evaluation⏱ 14 min read📅 Updated June 2026

Models are said to pass the Bar Exam and turn out spotless code — but on what basis do we call them "smart"? Go inside the lab and see the benchmarks, Turing tests and adversarial evals that researchers use in 2026 to put a number on machine intelligence.

◆知微•🧪 AI Evaluation · ⏱ 14 min read · June 23, 2026

又一场发布会,又一轮「前所未有的智能」——这个领域大致就是这个节奏。宣传里通常还会配上 SAT 百分位第 90、Python 代码毫无瑕疵、通过医师执照考试之类的话。作为用户,你完全有理由反问:这些结论是靠什么得出的?模型到底在思考,还是只是极擅长做题?

计算机科学里少有比这更棘手的问题。人类 IQ 测试经过一个世纪的打磨,AI 评测却没有这样的稳定基础——靶子一直在移动。下面我们掀开帘子,看看模型是在哪里被打分的:基准测试题库、对抗性攻击,以及那些判断 AI 是否真聪明的「体感检验」。

01第一道难关:对机器来说,「聪明」到底指什么

任何测试都以定义为前提,而给机器定义智能,会直接把你带进哲学。把整部百科全书背下来,算聪明吗?能写出一首好诗,却不知道 2+2=4,这算有智能吗?

变通办法是把智能拆开,一个领域一个领域地看:

  • 知识检索:能否准确调取历史、科学、法律等领域的事实?
  • 逻辑推理:能否分多步解出数学题,或从给定前提推出结论?
  • 编程与空间感知:能否写出跑得起来的软件,能否理解物体之间的空间关系?
  • 语言的细微之处:能否听懂讽刺、习语和复杂的指令?

这样分类之后,智能就能逐领域施测。最终得到的不是机器的一个「IQ 分数」,而更像一份覆盖多门学科的成绩单。

02经典图灵测试,以及研究者为什么不再用它

一提到 AI 测试,图灵测试迟早会被搬出来。它由艾伦·图灵(Alan Turing)在 1950 年提出,形式极简:评判者通过文字同时与一个人和一台机器交谈;如果无法有把握地分辨哪个是哪个,机器就算通过。

如今它的地位更像一座历史里程碑,而不是一件可用工具。要真正衡量机器智能,需要更严格的手段——可量化,而且不容易被「钻空子」。

03基准测试:今天真正在考的卷子

于是基准测试登台。它们是体量庞大的标准化数据集,包含横跨多个学科的数千道题目;新模型训练完成后,就用它们跑分。2026 年分量最重的几项如下:

基准测试考察内容难度
MMLU大规模多任务语言理解(Massive Multitask Language Understanding)——涵盖物理、法律、医学等 57 个学科本科 / 研究生水平
HumanEval代码生成:模型必须写出可运行的 Python,解出算法题。在职工程师水平
GSM8K小学数学:多步骤应用题,靠的是推理,而不只是算术。初中水平
ARC-Challenge高级推理:解题需要常识加背景知识。高中水平

成绩单该怎么读

评测从来不看单一数字。研究者看到的是把各项成绩绘成的一张雷达图。如果某模型在 MMLU(知识)上拿到 95%,而在 ARC(推理)上只有 40%,结论就是:它读过很多书,但批判性思维偏弱。

04红队测试:靠打碎它来检验智能

数学考得好,不说明它安全,也不说明它真的懂。红队测试填的就是这个空当。实验室请来专家——善意黑客、语言学家、领域专家——请他们刻意去「打碎」系统,或者诱使它暴露出自身的局限。

红队评估的流水线怎么走
  1. 1

    对抗性提示
    专家用复杂的逻辑谜题和「越狱」手法发问,看模型的推理在压力下是否还站得住。



    2

    边缘情形测试
    给视觉与多模态系统喂入极端、古怪的输入——思路类似我们在 AI 深度伪造是什么、如何识别 里分析问题素材的做法。



    3

    安全边界测绘
    找到模型「智能」崩塌、开始输出有害或无意义内容的确切临界点。



    4

    真实场景模拟
    推演模型失效会造成真实伤害的场景,预防 AI 被用于诈骗与欺诈 这类情况发生。

如果模型会被一个普通的逻辑谬误骗倒,不管基准分数多高,都会被记为推理能力薄弱。

05人工评估:也就是所谓的体感检验

数字永远说不清全部。模型可能解得出困难的物理方程,可讲解的文字却让人一头雾水、机械生硬,甚至隐隐带着居高临下的味道。人类偏好评估(Human Preference Evaluations)就是为了抓住这一点。

具体做法是:评判者看到两个不同模型针对同一提示词给出的回答,然后挑出在他看来更有帮助、更诚实、危害更小的那一个。这些判断随后回流到训练中——也就是 RLHF 技术——让系统学会人类在所谓「聪明」的回答里真正看重什么。

06AI 测试的漏洞:古德哈特定律

统计学里有一句名言叫古德哈特定律(Goodhart's Law):一旦某个指标变成了目标,它就不再是好的指标。今天没有比这更让 AI 评测头疼的事了。

各家实验室都在 MMLU 这类基准上争夺靠前的名次,针对这些具体试卷做优化的动力因此极大,随之而来的是两个熟悉的问题:

  • 数据污染:模型在基于海量网络抓取数据的训练中,无意间把基准题目「背」了下来。那 90% 反映的是它提前见过这套题,而不是智能。
  • 基准刷分:开发者可以做一些细微调整,去迎合某项基准的具体格式,而底层的推理能力和从前一样薄弱。

应对办法是转向动态评估(Dynamic Evals):题库不断变动、每周刷新,并严格保密。想追踪评测方法走向何方,可以看 本周 AI 研究动态,新的、未被污染的基准一出现就会被标出来。

07读者最常问的问题

实际中,AI 的智能是怎么测的?
主要靠三件工具:标准化基准测试——衡量知识用 MMLU,衡量编程用 HumanEval——用来搜寻安全弱点的对抗性红队测试,以及判断回答是否有帮助、是否自然的偏好评估(RLHF)。此外还有动态「评估」,实时考察推理能力,让背诵式记忆无法抬高分数。
研究者还在靠图灵测试判断智能吗?
如今学界大多认为它已被取代。它当年能检验机器能否在对话中冒充人类,却测不出推理、逻辑和解决问题的能力。多任务基准测试与对抗性测试已经接替了它的位置。
什么算 AI 基准测试?
基准测试指用来给模型表现打分的标准化数据集,或一套固定任务。MMLU 考察学术知识,GSM8K 考察数学应用题,ARC 考察逻辑推理。分数高,说明在该具体领域实力强。
AI 有可能在这些测试里作弊吗?
会,途径就是所谓的「基准刷分」和「数据污染」——模型在训练时无意间把测试题吸收了进来。解药是不断构建新的隐蔽题库,即「动态评估」,让被测量的是推理而不是记忆。
AI 评估里的「红队测试」是什么意思?
红队测试会请来善意黑客和领域专家刻意攻击模型——把它打坏、诱使它暴露缺陷,或者从它的安全护栏旁边绕过去。它衡量的是在充满敌意的条件下,这份智能还剩下多少稳健与安全。
◆

知微

我们关注的领域是 AI 前沿,力求把事实与科幻分开。本指南最近一次准确性审核是在 2026 年 6 月。对 AI 评测有疑问?联系我们的团队 或 了解我们的初衷。

Another launch event, another round of talk about "unprecedented intelligence" — that is the rhythm of this field. The pitch usually includes 90th percentile on the SATs, impeccable Python, a pass on the medical licensing exam. A sensible user is entitled to push back: on what evidence are such claims made, and is anything actually being thought? Or is this simply very skilled test-taking?

Few problems in computer science are messier. Human IQ testing has had a century of refinement behind it; AI evaluation has no such stability — the target keeps moving. What follows is a look behind the curtain at the places where models get graded: the human "vibe checks", the adversarial attacks, and the benchmark suites that together decide whether a system deserves to be called smart.

01The Hard Part: Defining What "Smart" Means for a Machine

Any test presupposes a definition, and defining intelligence for a machine drags you straight into philosophy. Memorising an entire encyclopedia — does that count as smart? Turning out a lovely poem while failing to grasp that 2+2=4 — is that intelligent?

The workaround is to split the concept apart, domain by domain:

  • Knowledge Retrieval: how reliably does it pull up facts from history, science and law?
  • Logical Reasoning: how does it handle maths that unfolds over several steps, or draw a conclusion from stated premises?
  • Coding & Spatial Awareness: does it produce software that runs, and does it grasp physical relationships?
  • Linguistic Nuance: how well does it handle sarcasm, idioms and complicated instructions?

Broken into categories like these, intelligence becomes testable domain by domain. What emerges is not one "IQ score" for the machine but something closer to a full report card spanning many subjects.

02The Turing Test, and Why Researchers Moved On

Raise the subject of AI testing and the Turing Test will come up before long. Alan Turing put it forward in 1950, and the setup is minimal: a judge exchanges text with one person and one machine; if the judge cannot tell them apart with confidence, the machine has passed.

Its standing nowadays is that of a historical landmark rather than a working instrument. Genuinely measuring machine intelligence called for something stricter — something you can quantify and that resists being "gamed".

03Benchmarks: Today's Actual Intelligence Exams

This is where benchmarks took over. They are enormous standardized datasets, thousands of questions spanning many disciplines, and a freshly trained model is run through them to produce a score. The ones that carry most weight in 2026:

BenchmarkWhat It MeasuresDifficulty
MMLUMassive Multitask Language Understanding — 57 subjects spanning physics, law and medicineCollege / Graduate
HumanEvalCode generation: the model has to produce working Python that solves algorithmic puzzles.Practising Engineer
GSM8KMulti-step word problems from grade-school maths, demanding deduction rather than mere arithmetic.Middle School
ARC-ChallengeAdvanced reasoning: solving these calls for common sense plus background knowledge.High School

Reading the Report Card

Evaluation never rests on a single number. The picture is a radar chart plotting performance across the whole set. A model at 95% on MMLU (knowledge) yet 40% on ARC (reasoning) tells researchers it has absorbed a great deal of schooling while falling short on critical thinking.

04Red Teaming: Probing Intelligence by Breaking It

A strong maths score says nothing about safety or genuine understanding, which is the gap red teaming fills. Experts are brought in — friendly hackers, linguists, domain specialists — and paid to "break" the system deliberately, or to lure it into exposing where its limits lie.

How the Red Teaming Pipeline Runs
  1. 1

    Adversarial Prompting
    Specialists throw intricate logic puzzles and "jailbreaks" at the model to see whether its reasoning holds up under strain.



    2

    Edge Case Testing
    Visual and multimodal systems get fed odd, extreme inputs — much the way flawed media is examined in our piece on what an AI deepfake is and how to spot one.



    3

    Safety Boundary Mapping
    Locating the precise threshold at which the model's "intelligence" collapses and harmful or incoherent output begins.



    4

    Real-World Simulation
    Running through scenarios in which a model failure causes genuine harm, heading off the kind of situation covered in scams and fraud built on AI.

A model that a commonplace logical fallacy can trip up will be marked down for weak reasoning, however lofty its benchmark numbers happen to be.

05Human Evaluation, a.k.a. the Vibe Check

Numbers never capture everything. A model can crack a difficult physics equation and then explain it in prose that is baffling, mechanical or faintly patronising. Human Preference Evaluations exist to catch exactly that.

Concretely: a rater sees two replies, produced by two different models for one prompt, and picks the one that strikes them as more helpful, more honest and less harmful. Those verdicts then feed back into training — the technique known as RLHF — teaching the system what people actually prize in an answer they would call smart.

06Where AI Testing Breaks Down: Goodhart's Law

Statistics has a well-known adage for this, called Goodhart's Law: once a metric is treated as the goal, it stops working as a metric. Nothing troubles AI evaluation more today.

Since labs compete for top placement on suites such as MMLU, the incentive to tune for those exact tests is enormous — and it produces two familiar problems:

  • Data Contamination: the model "memorises" benchmark questions inadvertently while training on a vast web crawl. That 90% reflects prior exposure to the exam, not intelligence.
  • Benchmark Hacking: developers can make small adjustments that suit a given benchmark's exact format, while the reasoning underneath stays exactly as weak as before.

The countermeasure is the move to Dynamic Evals: sets that shift constantly, get refreshed every week and stay under strict secrecy. If you want to track where evaluation methods are heading, this week in AI research will flag the arrival of new benchmarks that have not been contaminated.

07Questions Readers Ask Most

In practice, how is AI intelligence measured?
Three instruments do most of the work: standardized benchmarks — MMLU for knowledge, HumanEval for coding — adversarial red-teaming that seeks out safety weaknesses, and human preference evaluations (RLHF) that judge how helpful and natural a reply sounds. On top of that sit dynamic "evals", which exercise reasoning live so that memorisation cannot inflate the result.
Do researchers still rely on the Turing Test for intelligence?
Most researchers in the field now regard it as superseded. It once served to show whether a machine could pass for a human in conversation, but it cannot tell you about reasoning, logic or problem-solving. Multi-task benchmarks and adversarial testing have taken its place.
What counts as an AI benchmark?
Benchmark describes a standardized dataset, or a fixed task set, used to grade how a model performs. MMLU tests academic knowledge, GSM8K tests maths word problems, ARC tests logical reasoning. A high score means strength in that particular domain.
Is it possible for AI to cheat on these tests?
It can, through what are called "benchmark hacking" and "data contamination" — the model inadvertently absorbing the test questions while training. Fresh hidden test sets, known as "dynamic evals", are continuously built as the antidote, so that reasoning rather than recall is what gets measured.
What does "Red Teaming" mean in AI evaluation?
Red teaming brings in friendly hackers and subject experts to attack a model on purpose — breaking it, tricking it into showing its weaknesses, or slipping past its safety guardrails. What it measures is how robust and how safe the intelligence remains under hostile conditions.
◆

知微

Our beat is the frontier of AI, where we try to keep fact separate from science fiction. Accuracy checks on this guide were last carried out in June 2026. Questions about how AI gets evaluated? Reach our team or read more about what drives us.