RT-2 这类机器人基础模型浅释Robot Foundation Models Such as RT-2: A Plain Look

🧠 具身 AI⏱28 分钟阅读

网络规模的知识正在转化为具体动作。本文介绍 Google 的 RT-2 及其他视觉-语言-动作(VLA)模型如何让机器人获得涌现推理能力,重塑 2026 年的自主操作。

◆知微•🧠 具身 AI · ⏱28 分钟阅读 · 2026 年 9 月 16 日
🧠 Embodied AI⏱ 28 min read

Internet-scale knowledge is becoming physical action. See how Google's RT-2 and other Vision-Language-Action (VLA) models hand robots emergent reasoning and reshape autonomous manipulation in 2026.

◆知微•🧠 Embodied AI · ⏱ 28 min read · September 16, 2026

很长一段时间里,人工智能与机器人学各走各路。在数字世界里,AI 处理语言和图像得心应手;机器人则依赖预先写死的僵硬规则与物理世界打交道。如今这一格局已经反转——今天打造的机器人不再只是照本宣科,它们会理解。

「基础模型」这个说法,你或许在大型语言模型(LLM)的报道里听过。那么,RT-2 这样的机器人基础模型究竟是什么?人们又为何称它为具身 AI 的转折点?简单说,它是一个规模庞大、能力通用的 AI,在混合数据集上训练而成——既有互联网规模的文本与图像,也有真实机器人交互数据——能把人给出的高层指令转化为精确动作。

最具代表性的例子是 Google DeepMind 的 RT-2(Robotic Transformer 2)。它把机器人运动当作另一种可供预测的「语言」,于是 RT-2 及同类视觉-语言-动作(VLA)模型让机器拥有了曾被认为遥不可及的常识判断力:它们能认出陌生物体、理解抽象概念、适应从未到过的环境,而这一切都不需要重新手写指令。

01机器人 AI 之路:从单一任务到广泛能力

要理解 RT-2,先要看清传统机器人 AI 的局限。早期机器人靠「狭窄」的任务专用模型学习:想让机械臂拿起红方块,就得让它在成千上万个红方块样本上反复训练;一旦换成蓝方块、海绵或没见过的工具,它就彻底失灵。

这种做法从根本上无法规模化。真实世界里物体形态、光照条件、空间布局数以百万计,为每一种可能的场景手工标注训练数据,在经济上和实操上都走不通。

当研究者把语言领域的「基础模型」思路借过来,突破口随之出现。LLM 通过阅读整个互联网掌握语言的内在结构;机器人基础模型则借助海量而多样的数据,掌握物理交互的内在结构。RT-2 正是这一思路的集大成者,把网络丰富的语义与机器人对物理世界的扎根结合在一起。

02RT-2 详解:机器人 transformer

RT-2(Robotic Transformer 2)由 Google DeepMind 与合作机构共同打造,是一个视觉-语言-动作(VLA)模型。它延续了 RT-1 的成功,却带来一处决定性的架构变化:在机器人轨迹数据之外,还同时加入来自网络的大规模视觉-语言数据进行联合训练。

100+
支持的机器人类型
Open X-Embodiment 数据集
2x
在陌生任务上的表现
与前代 RT-1 对比
50%
抽象概念成功率
涌现推理测试

RT-2 真正的威力在于知识迁移。视觉-语言预训练相当于让它「读」过互联网,因此它知道 croissant 是什么、已经灭绝的动物(比如玩具恐龙)长什么样,也知道打翻的液体需要擦拭。收到指令时,它不只是匹配视觉模式,而是从庞大的语义知识库出发进行推理。

03视觉-语言-动作(VLA)架构

简单看看架构,就能明白 RT-2 是怎么运行的。传统机器人流程是割裂的:一个模型识别物体,另一个规划路径,第三个再驱动电机;VLA 则把这些环节收进同一张端到端网络。

把机器人控制当作序列建模问题后,RT-2 便能搭上造就 LLM 的同一套缩放规律:接触的数据越多样,肢体行为就越稳健、越通用。

04涌现推理:顿悟时刻

最能体现 RT-2 本事的,是它的「涌现推理」。测试中,研究者给出的指令需要多步逻辑推导,而不只是认出物体。

桌上摆着一堆东西,再让机器人「拿起已经灭绝的动物」,传统机器人因为没有这个类别而无从下手。RT-2 则调用网络规模的训练成果,把「已经灭绝的动物」对应到桌上的玩具恐龙,成功把它抓起;让它「拿起对付头痛最管用的东西」时,它也准确挑出了一瓶布洛芬。

这些都不是事先写好的行为,而是真正的零样本泛化——模型把语言知识和世界知识用到了一个从未受过专门训练的操作任务上。这延续了训练 AI 机器人抓取物体的进展,却把它从底层运动控制提升到了高层语义理解。

05机器人基础模型的现实应用

目前 RT-2 主要还是研究层面的里程碑,但 VLA 的走向预示着它将在众多行业产生深远的实际用途:

1. 柔性制造与物流

传统工厂机器人被螺栓固定在地面,只负责一道工序。基础模型则有望带来「通用型」机器人:一进仓库,只要告诉它要找什么,就能立刻分拣形状不规则、从未见过的包裹,随库存变化而调整,省去数周的重新编程。

2. 家务与养老协助

家居环境高度非结构化、充满意外。搭载 RT-2 式模型的机器人可以听懂「把厨房洒的东西擦干净」「把我的老花镜拿来」,带着窄模型不具备的情境意识,在有人居住的杂乱空间里穿行。

3. 危险环境探索

在灾害救援或太空探索中,通信延迟让人难以实时遥控。给模型一个高层目标——「在废墟里搜寻幸存者」或「沿那道山脊采集岩石样本」——它就能自己推演出所需的复杂肢体步骤,只有遇到真正无法逾越的边缘情况时,才升级为机器人遥操作。

06当前挑战与局限

尽管热度很高,这类模型并不是万能药。要走向普及,还必须跨过若干重大障碍:

挑战具体说明现有缓解办法
数据稀缺高质量、多样化的机器人交互数据,比互联网文本稀有得多,差距以指数计。Open X-Embodiment 等协作数据集,以及合成数据生成。
算力成本训练和运行庞大的 VLA 模型需要巨量算力,限制了边缘部署。模型蒸馏、量化,以及专用 AI 加速器。
安全与对齐能做抽象推理的模型,也可能幻觉出危险的肢体动作。严格的沙箱测试、奖励建模,以及人在回路中的监督。
仿真到现实的鸿沟正如关于人形机器人最大挑战的讨论所言,把仿真中学到的东西迁移到真机上依然困难。域随机化,以及在真机上的大量微调。
对抗性漏洞对抗性贴片或被操纵的输入可能骗过视觉系统。建立可靠的 AI 深度伪造识别与传感器校验流程,确保数据可信。

安全挑战格外尖锐:LLM 幻觉时吐出的是无意义文字;机器人基础模型一旦幻觉,则可能用力过猛、失手掉落危险品,或做出无法预判的动作。因此研究者把「物理对齐」——让机器人的动作始终安全、有益——视为首要任务。

07具身 AI 的未来

RT-2 不是终点,而是一块奠基性的踏脚石。研究界已在推进 Open X-Embodiment——一场大规模协作,力图用全球数十种机器人贡献的数据,训练同一个通用机器人「大脑」。

在不远的将来,可以期待以下方向:

  • 多模态基础模型:在视觉和语言之外,把触觉、声音和本体感觉也纳入进来,形成对世界完整的感官理解。
  • 持续学习:机器人从现实世界的每一次交互中学习,增量更新基础模型,而不发生灾难性遗忘。
  • 机器人技术的民主化:正如云 API 让任何软件开发者都能用上先进 AI,由基础模型驱动的「机器人即服务」(RaaS)API 也将让任何公司无需从零训练,就能调用先进的机器人能力。

真正的通用型机器人——叠完衣服转而去装配电子产品、再到医院帮忙——这一梦想已经不再局限于科幻。RT-2 这类机器人基础模型已经给出了把它变为现实的架构蓝图。

08常见问题

RT-2 这样的机器人基础模型到底是什么?
RT-2(Robotic Transformer 2)是一种大规模 AI,同时在两类海量数据上训练:互联网文本与图像,以及机器人交互记录。传统模型只为单一任务服务,而 RT-2 的视觉-语言-动作(VLA)结构使其能借助网络规模的语义知识,把人类给出的高层指令转化为精确的肢体动作。
RT-2 与以往的机器人 AI 有何不同?
以往的机器人 AI 依赖狭窄的任务专用数据(比如只练习拿起红方块)。RT-2 则在机器人数据与网络规模的视觉-语言数据上联合训练,从而表现出「涌现推理」——理解陌生物体、抽象概念和从未被明确教过的复杂指令。
视觉-语言-动作(VLA)模型是什么?
视觉-语言-动作(VLA)模型是这样一种架构:它接收摄像头画面和文字指令,直接输出动作 token(即机器人电机指令)。把动作当作另一种可供预测的「语言」,VLA 模型便能以一张统一网络,在众多机器人和任务之间泛化。
机器人基础模型会被入侵或欺骗吗?
会。没有哪种 AI 系统能完全免疫对抗攻击;被篡改的画面可能让机器人误认物体或做出不安全的动作。正因如此,需要类似 AI 深度伪造检测所采用的严密防护,来核验物理 AI 系统中传感器数据的完整性。
消费者什么时候能用上机器人基础模型?
目前它们主要处于前沿研究和企业试点阶段。如果发展速度延续,经过缩小和优化的版本有望在 2020 年代后期开始进入家用辅助机器人等消费级产品,前提是成本和安全问题得到解决。
◆

知微

我们持续追踪全球 AI 与机器人动向,帮你读懂塑造具身智能未来的技术。本文已于 2026 年 9 月完成准确性审核。有问题?欢迎联系我们的团队,或了解我们的使命。

For many years, artificial intelligence and robotics ran on separate lines. Inside the digital world, AI handled language and images with skill; robots, by contrast, depended on inflexible rules coded in advance to touch the physical world. That order has now flipped. The robots being built today do more than run through scripts—they understand.

The phrase foundation model may already ring a bell from coverage of large language models (LLMs). What exactly is a robot foundation model in the style of RT-2, and why do people call it a turning point for embodied AI? Put plainly, it is a very large, broadly capable AI trained on mixed datasets—text and images at internet scale alongside data from real robotic interaction—that turns a person's high-level request into exact movement.

The leading case in point is Google DeepMind's RT-2 (Robotic Transformer 2). By treating robot motion as yet another kind of language worth predicting, RT-2 and fellow Vision-Language-Action (VLA) models give machines a brand of common-sense judgment once assumed out of reach: they spot unfamiliar objects, grasp abstract ideas, and adjust to settings they have never met, all without fresh hand-coded instructions.

01Robot AI's Path: From Single Tasks to Broad Skill

RT-2 only makes sense against the backdrop of what older robotic AI could not do. Earlier machines learned through narrow, single-job models. Teach an arm to lift a red block and that was all it knew, trained on thousands upon thousands of red-block lifts. Swap in a blue block, a sponge, or an unfamiliar tool, and performance collapsed.

That strategy cannot scale, at root. The physical world offers countless object shapes, lighting states, and spatial layouts; labeling training examples by hand for each conceivable situation is neither affordable nor feasible.

The door opened once researchers borrowed the foundation-model idea from language work. An LLM absorbs the shape of language by reading across the internet; in much the same way, a robot foundation model absorbs the shape of physical interaction from broad, massive data. RT-2 is that idea fully realized, joining the web's semantic depth to robotics' grounding in the physical.

02RT-2 Explained: The Robotic Transformer

Built by Google DeepMind together with partner institutions, RT-2 (Robotic Transformer 2) is a Vision-Language-Action (VLA) model. It carries forward RT-1's wins with one decisive change: training now mixes robot trajectory data and large-scale vision-language material drawn from the web.

100+
Robot types it supports
The Open X-Embodiment dataset
2x
Results on unseen tasks
Compared with the earlier RT-1
50%
Success with abstract ideas
Tests of emergent reasoning

RT-2's real power is knowledge transfer. Its vision-language pre-training amounts to reading the internet, so it recognizes what a croissant is, how an extinct animal such as a toy dinosaur appears, and that a spill calls for wiping. A command prompts more than pattern matching; the machine reasons from a deep store of semantic knowledge.

03The Vision-Language-Action (VLA) Blueprint

A short tour of the architecture shows how RT-2 operates. Classic robot stacks are split apart: a detection model here, a path planner there, a third model driving the motors. VLA folds those pieces into one end-to-end network.

Cast robot control as sequence modeling and RT-2 can ride the same scaling laws that powered large language models: wider, more varied data yields steadier, more general physical behavior.

04Emergent Reasoning: The Sudden Leap

Nothing shows RT-2's reach like its emergent reasoning. In trials, researchers issued orders that demanded chains of logic rather than mere object recognition.

Given a spread of objects and the order to pick up the extinct animal, an ordinary robot stalls for want of such a category. RT-2 instead drew on its web-scale training, linked extinct animal to the toy dinosaur lying on the table, and grasped it. Told to grab what works best for a headache, it likewise singled out a bottle of ibuprofen.

None of that was coded in advance; this is true zero-shot generalization. Language knowledge and world knowledge are applied to a manipulation task the model was never trained on. The result extends progress in training AI robots to grip, lifting it from basic motor control toward high-level semantics.

05Where Robot Foundation Models Meet the Real World

RT-2 remains, for now, chiefly a research landmark, yet the direction of VLA points to deep practical uses in field after field:

1. Bendable Manufacturing and Logistics

Conventional factory robots sit bolted down and run one job. Foundation models promise general-purpose machines that can walk into a warehouse and start sorting odd, never-seen packages as soon as someone names the target, shifting with new stock instead of spending weeks in reprogramming.

2. Help at Home and in Elder Care

Living spaces are unstructured and full of surprises. Running an RT-2-style model, a robot could act on clean up the kitchen spill or fetch my reading glasses, threading through household clutter with contextual awareness narrow models cannot match.

3. Probing Hazardous Settings

Disaster zones and space missions impose communication lag that blocks direct human steering. Handed a broad goal—search the rubble for survivors, or gather rock samples along that ridge—a well-grounded model can work out the physical steps alone, falling back on robotic teleoperation only for an edge case it truly cannot clear.

06Hurdles Today: What Still Holds Them Back

Enthusiasm aside, these models are no cure-all. A number of serious barriers stand between them and everyday use:

ObstacleWhat it involvesHow teams address it
Shortage of dataCompared with web text, rich and varied robotic interaction data is exponentially harder to find.Shared sets such as Open X-Embodiment, plus synthetic-data generation.
The bill for computeEnormous computing power is needed to train and run large VLA models, which hampers deployment at the edge.Distillation, quantization, and purpose-built AI accelerators.
Safety and alignmentAbstract reasoning can also lead a model to hallucinate a dangerous physical move.Sealed sandbox trials, a modeled reward signal, and reviewers kept inside the loop.
The sim-to-real divideAs raised in coverage of humanoid robotics' hardest problem, moving simulated learning onto real machines stays troublesome.Domain randomization and extensive fine-tuning on hardware.
Weakness to adversarial tricksPatches or manipulated inputs can mislead a vision-driven machine.Reliable pipelines for spotting AI deepfakes and validating sensor input, keeping data trustworthy.

The safety problem carries unusual weight. A hallucinating LLM prints nonsense; a hallucinating robot may exert too much force, drop something hazardous, or lurch without warning. Researchers therefore put physical alignment—movements that stay safe and useful—above all else.

07What Comes Next for Embodied AI

RT-2 marks a starting block rather than a finish line. Work is already turning toward Open X-Embodiment, a broad cooperative push to build one universal robot brain from data contributed by dozens of machine types worldwide.

Near-term expectations include the following:

  • Multimodal foundation models: beyond sight and language, weaving in touch, sound, and proprioception for a full sensory read on the world.
  • Continuous learning: machines that learn from each real-world encounter and update their base model gradually, without catastrophic forgetting.
  • Wider access to robotics: cloud APIs put advanced AI within any developer's reach; likewise, foundation-model Robotics-as-a-Service (RaaS) APIs will let any firm call up strong robotic skill without training a model from zero.

The vision of a genuine general-purpose helper—folding laundry, then building electronics, then assisting in a hospital—has escaped the pages of science fiction. RT-2-style foundation models supply the architectural map to build it.

08Common Questions

What exactly is an RT-2-style robot foundation model?
RT-2 (Robotic Transformer 2) represents a large-scale AI trained across two kinds of massive data: internet text and images, and robotic interaction logs. Older models served one task; RT-2's Vision-Language-Action (VLA) structure lets it convert high-level human requests into exact movement by drawing on web-scale semantics.
How does RT-2 stand apart from earlier robot AI?
Earlier robot AI learned from narrow, single-purpose data—red-block lifts and nothing else. RT-2 trains jointly on robot and web-scale vision-language material, which unlocks emergent reasoning: unseen objects, abstract notions, and involved commands it was never explicitly taught to perform physically.
What does a Vision-Language-Action (VLA) model do?
A Vision-Language-Action (VLA) model is a design that takes in pictures from cameras and written requests and returns action tokens—motor commands—directly. Treating motion as another language to predict lets one unified network generalize across many robots and jobs.
Can robot foundation models be hacked or fooled?
They can be. No AI system is immune to adversarial attack; a doctored image might make the machine misread an object or choose an unsafe move. That is why strong safeguards akin to those behind AI deepfake detection are needed to verify sensor data in physical AI.
When will consumers get to use robot foundation models?
Today these systems mostly live in advanced research and corporate pilots. If progress keeps its pace, scaled-down, optimized versions could power consumer products such as sophisticated home assistants starting in the late 2020s, assuming cost and safety get resolved.
◆

知微

Our beat is AI and robotics around the world, tracked so the forces behind embodied intelligence make sense to you. Facts reviewed in September 2026. Questions? Write to our team or read what we're about.