斯坦福把论文做成了 AI Agent,这事上了 Nature 正刊Stanford Turned a Paper Into an AI Agent, and It Landed in Nature
斯坦福团队做了个叫 Paper2Agent 的工具:把论文连同代码仓库,自动打包成 AI 助手能直接调用的接口,成果上了 Nature 正刊。它怎么干的、花了多少钱、替谁省了事。
A Stanford team built a tool called Paper2Agent: it automatically packages a paper, together with its code repository, into an interface an AI assistant can call directly — and the work landed in Nature. How it does it, what it costs, and whose work it saves.

摘要
斯坦福团队做了个叫 Paper2Agent 的工具:把论文连同代码仓库,自动打包成 AI 助手能直接调用的接口,成果上了 Nature 正刊。它怎么干的、花了多少钱、替谁省了事。
正文
先说个场景
你在 arXiv 上刷到一篇论文,方法看着正对你的路子,想把它的流程套到自己的数据上。
接下来大概会是这么两个小时:clone 仓库、建环境、装依赖、跟版本冲突打架、跑教程、报错、改参数、再报错。等终于跑通,你已经忘了自己一开始想验证什么。
斯坦福的一群人显然也被这件事折磨过。他们做了个工具叫 Paper2Agent,2026 年 9 月 16 日上了 Nature 正刊,论文标题直译过来是《把研究论文重新想象成可交互、可靠的 AI Agent》。
说人话就是:把一篇论文连同它的代码仓库,全自动打包成一个 AI 助手能直接调用的接口。以后你不用读教程、跑代码,直接用自然语言跟它说话就行。
作者给它安了个挺好听的称呼——一个虚拟的通讯作者。
它到底交出来一个什么东西
输入端是一篇论文加它的代码仓库,输出端是一个 MCP server。
MCP 全称 Model Context Protocol,是个让各家 AI 助手统一调用外部工具的标准协议。你可以把它理解成一种"插头规格"——插头做成这个规格,Claude、Codex 之类的 agent 都能往上插。
这个 server 往外提供三样东西。
- MCP tools:把论文里的方法包装成可以直接调用的函数
- MCP resources:手稿、代码链接、数据集、图
- MCP prompts:把多步骤流程整个写进去,比如 Scanpy 那套正确的预处理顺序
已经有三个现成的公开挂着,名字叫 AlphaGenome、Scanpy、TISSUE。
六步流水线,加一道及格线
流水线搭在 Claude Code 的 agent SDK 上,一个中心编排器指挥几个专职小 agent,依次干六件事。
- 找到并下载代码仓库
- 环境管理器搭一个隔离的虚拟环境
- 教程扫描器给能用的教程建索引
- 教程执行器把教程端到端跑一遍,记下参考输出
- 工具抽取器把教程转成参数化的 MCP 工具,交给测试验证器
- 通过验证的工具,由编排器组装成一个 server
中间那道及格线值得单独说一句。一个工具要同时满足两条才算过:该出现的文件出现了,而且数值和参考输出的误差在 3% 以内;如果涉及图,还要过一道感知哈希比对,汉明距离得小于 20。每个函数最多给 6 次机会,反复不过的直接剔掉,不进最终的 server。
整套东西用的是 Claude Sonnet 4。
装好之后,用法就是 Claude Code 里一个 /mcp 条目的事:

账单
养一个 agent 要花多少钱?这组数字挺有诚意,是作者自己报出来的。
AlphaGenome 这个案例,22 个工具,约 45 分钟,14 美元,22 个全部通过验证,全程没有人插手。
Scanpy 那个,7 个工具,约 45 分钟,13 美元,还在 4 个公开数据集上和人类研究员对上了细胞数、基因数、top marker 基因。
规模测试更狠一点:100 篇 bioRxiv 计算生物学论文,74 篇成功 agent 化,599 个候选工具里 593 个过了验证,全程没有人手工收拾。
单次查询 0.20 美元、1.6 分钟;对照组 0.38 美元、4.3 分钟。
准确率上,教程衍生查询 98.7%(Claude 加仓库访问权限 82.7%,Biomni 37.3%);全新查询 100.0%(78.7%、56.0%);开放式查询 82.7%(56.7%、72.2%)。数据是 5 次运行、两位人类专家打分,评分者一致性 96.7%。基线换成更强的 Claude Opus 4.6 之后,优势依然成立。
300 道题的综合得分 91.2%,对照的两组是 80.3% 和 86.3%。10 篇非生物学论文(含 TabPFN、SAM 2、SAELens)在 42 个执行任务上 98.1%。26 篇数据类论文,resource 层得分 89.0%,浏览器方案 82.0%,还便宜 34 倍、快 15 倍。
置换测试里,超出范围的查询 100% 被拒;往里塞依赖错误、文件路径错误、拼写错误、废弃 API,它都能自己缓过来。
这组数字里有个挺实在的东西:100 篇里有 74 篇能成。剩下那 26 篇不是它不行,是那些论文本身没留下能跑的代码和教程。它的适用面,由别人写下的工程细节决定。
它真的被拿去做了点研究
有个案例挺有意思。研究者把三个 agent 串起来:AlphaGenome、一个 MPRA 偶联的 scCRISPRi 筛选(用大规模实验测调控元件的作用)、还有一份 CD4+ T 细胞的 Perturb-seq 数据集。
AlphaGenome 在银屑病相关位点 rs887314 上把 GPR137 标了出来,RNA-seq 分位得分 0.997。AI co-scientist 提了 10 个验证策略,研究者挑了 signature correlation 那条。结果只有 GPR137 敲低能匹配上 CRE 扰动特征,而且是在刺激条件下——Stim8hr 时 Spearman 0.613,Stim48hr 时 0.630,另外三个候选(含 BAD)没有显著相关。
第二条线把 AlphaGenome 和一个 ADHD 的全基因组关联研究连起来,在 209 个候选位点里提名了 rs1626703。这条目前还停在假设上,等实验来验。

一个必须写清楚的地方
论文里有个 LDL 胆固醇变异 chr1:109274968:G>T 的例子。agent 重新分析之后,把 SORT1 排成了最可能的因果基因;而原 AlphaGenome 论文强调的是 CELSR2 和 PSRC1。GTEx 数据库显示这三个基因在肝脏都有显著 eQTL。
团队自己拿这个例子说明的是:这类位点上判断因果基因本身就有多难。它给出的是一个不同的排序,不是纠正了原论文。
所以它到底替谁省了事
到这里结论其实不复杂。
Paper2Agent 干的是件很明确的工程活:把论文里重复的、机械的那部分调用成本自动化掉。它抬高的是复现效率的地板,不是科研判断的天花板。
真正决定你能不能把它用起来的,是你手里那篇论文里,有没有一份写得足够清楚的代码和教程。
数据来源与说明
本文数据与结论均来自该研究及其公开材料。论文 2026 年 9 月 16 日上线 Nature 正刊,标题为 Reimagining research papers as interactive and reliable AI agents,DOI 编号 10.1038/s41586-026-11044-y。预印本 2025 年 9 月 8 日上传 arXiv,编号 2509.06917,从预印本到正刊上线间隔约一年。项目以 MIT 许可发布,作为一个 skill 装进 Claude Code 或 Codex 使用。作者共五人:Jiacheng Miao、Joe R. Davis、Yaohui Zhang、Jonathan K. Pritchard、James Zou。文中三方对比数据按原口径完整给出,未做取舍。本文仅作技术梳理,与斯坦福、Nature 及任何机构不存在隶属、授权或背书关系。
Abstract
A Stanford team built a tool called Paper2Agent: it automatically packages a paper, together with its code repository, into an interface an AI assistant can call directly — and the work landed in Nature. How it does it, what it costs, and whose work it saves.
Body
Let's start with a scenario
You're scrolling arXiv and a paper's method looks exactly like what you need. You want to apply its pipeline to your own data.
What follows is roughly two hours of this: clone the repo, build the environment, install dependencies, fight version conflicts, run the tutorial, hit an error, tweak a parameter, hit another error. By the time it finally runs, you've forgotten what you originally wanted to verify.
A group at Stanford has clearly been tormented by the same thing. They built a tool called Paper2Agent, which landed in Nature on September 16, 2026. The paper's title translates roughly to "Reimagining research papers as interactive and reliable AI agents."
In plain terms: it takes a paper plus its code repository and automatically packages the whole thing into an interface an AI assistant can call directly. From then on, you don't read tutorials or run code — you just talk to it in natural language.
The authors gave it a rather nice label: a virtual corresponding author.
So what exactly does it hand you
The input is a paper plus its code repository. The output is an MCP server.
MCP stands for Model Context Protocol, a standard that lets AI assistants from different vendors call external tools in a uniform way. Think of it as a plug specification — build the plug to this spec, and agents like Claude or Codex can all plug into it.
The server exposes three kinds of things.
- MCP tools: the paper's methods wrapped as functions you can call directly
- MCP resources: the manuscript, code links, datasets, figures
- MCP prompts: entire multi-step workflows written in, such as the correct Scanpy preprocessing order
Three are already publicly available, named AlphaGenome, Scanpy, and TISSUE.

Six steps, plus a pass/fail line
The pipeline is built on Claude Code's agent SDK. A central orchestrator directs several specialist sub-agents through six jobs in order.
- Find and download the code repository
- An environment manager builds an isolated virtual environment
- A tutorial scanner indexes the tutorials that actually work
- A tutorial executor runs the tutorials end to end and records reference outputs
- A tool extractor converts the tutorials into parameterized MCP tools and hands them to a test validator
- Tools that pass validation are assembled by the orchestrator into a server
That pass/fail line in the middle deserves its own sentence. A tool has to satisfy two conditions at once to pass: the files that should appear do appear, and its values stay within 3% of the reference output; if figures are involved, it also has to pass a perceptual hash comparison with a Hamming distance under 20. Each function gets at most 6 attempts, and anything that keeps failing is dropped outright and never makes it into the final server.
The whole thing runs on Claude Sonnet 4.
Once installed, using it is a matter of one /mcp entry in Claude Code:

The bill
How much does it cost to run an agent? This set of numbers is refreshingly candid — the authors reported them themselves.
For the AlphaGenome case: 22 tools, about 45 minutes, $14, all 22 passing validation, with no human intervention at any point.
For Scanpy: 7 tools, about 45 minutes, $13, and it matched human researchers on cell counts, gene counts, and top marker genes across 4 public datasets.
The scale test is more aggressive: across 100 bioRxiv computational biology papers, 74 were successfully agentified, and 593 of 599 candidate tools passed validation, with no manual cleanup at any point.
A single query costs $0.20 and takes 1.6 minutes; the comparison condition costs $0.38 and takes 4.3 minutes.
On accuracy, tutorial-derived queries came in at 98.7% (Claude with repository access: 82.7%; Biomni: 37.3%); novel queries at 100.0% (78.7%, 56.0%); open-ended queries at 82.7% (56.7%, 72.2%). The data comes from 5 runs scored by two human experts, with 96.7% inter-rater agreement. When the baseline is swapped for the stronger Claude Opus 4.6, the advantage still holds.
Across 300 questions, the overall score was 91.2%, against 80.3% and 86.3% for the two comparison groups. Ten non-biology papers (including TabPFN, SAM 2, and SAELens) scored 98.1% on 42 execution tasks. Across 26 data-oriented papers, the resource layer scored 89.0% versus 82.0% for the browser approach — while being 34 times cheaper and 15 times faster.
In permutation tests, out-of-scope queries were rejected 100% of the time; when dependency errors, bad file paths, typos, and deprecated APIs were injected, it recovered on its own.
There's something quite grounded in these numbers: 74 out of 100 papers worked. The remaining 26 didn't fail because the tool can't do it — those papers simply didn't leave behind runnable code and tutorials. How broadly it applies is determined by the engineering details other people wrote down.
It really did get used for some research
One case is quite interesting. Researchers chained three agents together: AlphaGenome, an MPRA-coupled scCRISPRi screen (large-scale experiments measuring the effect of regulatory elements), and a Perturb-seq dataset from CD4+ T cells.
AlphaGenome flagged GPR137 at the psoriasis-associated locus rs887314, with an RNA-seq quantile score of 0.997. The AI co-scientist proposed 10 validation strategies, and the researchers picked the signature correlation one. The result: only GPR137 knockdown matched the CRE perturbation signature, and only under stimulated conditions — Spearman 0.613 at Stim8hr and 0.630 at Stim48hr, while the other three candidates (including BAD) showed no significant correlation.
A second thread connected AlphaGenome to a genome-wide association study of ADHD and nominated rs1626703 among 209 candidate loci. That thread remains a hypothesis, waiting on experiments to test it.

One point that has to be spelled out
The paper includes an example involving the LDL cholesterol variant chr1:109274968:G>T. After reanalyzing it, the agent ranked SORT1 as the most likely causal gene, whereas the original AlphaGenome paper emphasized CELSR2 and PSRC1. The GTEx database shows all three genes have significant eQTLs in liver.
What the team itself uses this example to illustrate is this: judging the causal gene at loci like this is inherently hard. It produced a different ranking; it did not correct the original paper.
So whose work does it actually save
By this point the conclusion isn't complicated.
Paper2Agent does a very specific piece of engineering work: it automates away the repetitive, mechanical part of the cost of invoking what's in a paper. It raises the floor of reproduction efficiency, not the ceiling of scientific judgment.
What really determines whether you can put it to use is whether the paper in your hands comes with code and tutorials written clearly enough.
Data sources and notes
The data and conclusions in this article all come from the study and its public materials. The paper went live in Nature on September 16, 2026, under the title Reimagining research papers as interactive and reliable AI agents, DOI 10.1038/s41586-026-11044-y. The preprint was uploaded to arXiv on September 8, 2025, as 2509.06917 — roughly a year between the preprint and the journal publication. The project is released under the MIT license and installs into Claude Code or Codex as a skill. There are five authors: Jiacheng Miao, Joe R. Davis, Yaohui Zhang, Jonathan K. Pritchard, and James Zou. The three-way comparison data in this article is given in full according to the original reporting, with nothing omitted. This article is a technical summary only, and has no affiliation, authorization, or endorsement relationship with Stanford, Nature, or any institution.
Figure Notes
| No. | File | Insertion point | Description |
|---|---|---|---|
| ------ | ------ | --------- | --------- |
| 1 | fig1-pipeline.png | After "So what exactly does it hand you" | Original paper Figure 1: the top half shows a paper converted into an MCP server and connected to any agent; the bottom half shows the six-step pipeline |
| 2 | mcp-connected.png | After "Six steps, plus a pass/fail line" | Official project screenshot: the /mcp list in Claude Code showing alphagenome connected |
| 3 | fig2-alphagenome.png | After "It really did get used for some research" | Original paper Figure 2: the full AlphaGenome case, including the SORT1 conclusion and validation accuracy |
| 4 | overview-card.jpg | Optional (insert at the top if length requires) | A media-produced key-points card with high information density, good for a "get it at a glance" opening |
Cover: reuse the already-generated cover-final.png (official logo overlaid with the Chinese headline).