The Paper That Answers Back: Stanford Turned 100 Research Papers into Agents
A research paper is a terrible interface to research.
Read the AlphaGenome paper — the one about predicting variant effects — and try to _use_ it. You will clone a repo, fight dependency versions for an afternoon, guess which notebook cell holds the actual method, and hand-translate the figures back into numbers. The paper describes the science. It refuses to _do_ the science.
A Stanford team — Jiacheng Miao, Joe R. Davis, Yaohui Zhang, Jonathan K. Pritchard, and James Zou — published the fix in **Nature on September 16, 2026** : **Paper2Agent** , an automated pipeline that converts a research paper into an MCP agent that reruns the paper's own methods on demand. You ask in English. It calls verified tools. You get the answer _and the code that produced it_.
The headline run: **100 computational biology papers in, 74 converted into working agents, 599 tools extracted, 593 passing automated validation.** And the agents beat the obvious alternative — handing Claude the same repos directly — **91.2% to 82.7%** on 300 benchmark questions.
## ELI5: the cookbook that cooks
A paper is a cookbook: ingredients listed, steps described, a photo of the finished dish. Every reader rebuilds the kitchen from scratch — different ovens, missing spices, a step the author forgot to write down. Half the time dinner fails and nobody knows why.
Paper2Agent turns the cookbook into a **chef you can talk to**. "Make me the variant-effect prediction for this mutation." The chef doesn't summarize the recipe — it _runs the kitchen_ , hands you the dish, and shows you the exact steps it followed, verified against the original photo.
The key insight: an LLM that has _read_ a repo is a worse scientist than an LLM that can _call_ the repo's verified functions. Reading is fuzzy. Calling is exact.
## How it works: six stages, zero humans
Paper2Agent takes a paper and does the whole miserable setup job itself:
1. **Locate the codebase** — find the repo, the notebooks, the docs linked from the paper.
2. **Build an isolated environment** — a sandbox where dependencies get installed and pinned.
3. **Scan the tutorials** — harvest the notebooks and docs that show the methods in action.
4. **Execute and capture** — run everything; save the true outputs as ground truth.
5. **Extract MCP tools** — wrap the runnable pieces as parameterized tools with JSON schemas.
6. **Assemble the MCP server** — a Model Context Protocol server any chat agent can call in natural language.
The validation gate between stages 5 and 6 is the whole ballgame: **numeric outputs must match the paper's own results within 3%** , generated figures are compared by **perceptual hash** , the tools are **locked** so the model can't improvise new behavior, and adversarial testing showed a **100% correct rejection rate** on out-of-scope queries. A tool that can't reproduce the paper's own outputs never ships.
## The numbers (SOTA section)
The scale run is what makes this a Nature paper and not a demo:
* **100 papers → 74 agents.** The 26 failures were blocked by missing code, unavailable data, or environments that wouldn't build. James Zou's framing: _"Agentification itself is a useful certificate that says, 'This work is relatively complete.'"_ The pipeline doubles as a reproducibility audit.
* **599 tools proposed → 593 validated (99.0%).** Six tools failed the reproduce-the-paper gate — exactly the gate working as designed.
* **91.2 ± 1.6% vs 82.7 ± 3.4%** on 300 tutorial-derived questions across all 74 agents. Same knowledge, same model family — the difference is _packaging_ : callable tools beat a pile of files.
* **The AlphaGenome agent specifically:** 22 tools built in ~45 minutes for ~$14 of compute. **98.7 ± 1.3%** on 15 tutorial queries, **100%** on 15 novel queries the authors wrote themselves (not in the tutorials).
* **Beyond biology:** 42 execution tasks from 10 non-biology computational papers hit **98.1 ± 0.8%** , with median runtime **1.9× faster** than Claude-plus-repo and **3.1× faster** than Biomni.
* **Economics:** $14 and ~45 minutes to agentify a paper. Per query: **$0.20 vs $0.38** baseline, **1.6 minutes vs 4.3 minutes** latency.
## The part that matters most: agents that team up
One agent answering questions is a better search engine. _Three agents collaborating_ is something new.
The authors built agents from three unrelated papers — an AlphaGenome variant-effect agent, an MPRA-coupled scCRISPRi screening agent, and a CD4+ T cell Perturb-seq agent — and let them chain computational predictions with experimental screens. Together they prioritized **GPR137 as a probable causal gene at the psoriasis-associated locus rs887314**. None of the three papers made that connection alone.
A separate Stanford Medicine demo: two agents from unrelated studies (a mutation-prediction tool and an ADHD genome-wide association study) surfaced a previously unreported variant near **MPHOSPH9** tied to increased ADHD risk. Zou's vision: _"millions of paper agents finding overlapping work at scale"_ — pairs of studies whose authors would otherwise never stumble into each other. He is blunt about the credit question: discoveries still get attributed back to the original papers and human authors.
## The limits, stated plainly
The paper doesn't soft-pedal them, so neither will I:
1. **26% of papers wouldn't convert.** Missing code, missing data, unbuildable environments. That's a reproducibility scandal wearing a success metric — and the authors know it.
2. **Humans still own the science.** Open-ended hypothesis selection and evidence evaluation stay human-in-the-loop. The agent reruns and applies methods; it doesn't decide what matters.
3. **Every agent rots.** MCP servers built this way need upkeep as upstream dependencies drift. A paper-agent is a living artifact with maintenance costs, not a PDF you file away.
4. **The "novel query" caveat.** The 100%-on-novel-queries result used 15 questions the _authors_ wrote. Independent benchmarks will be the real test.
Outside voices are interested but hedged — the right posture. Olivier Elemento (Weill Cornell): _"a real advance in terms of how we think about the publication process, with AI at the center."_ Dongping Chen (U. Maryland) called executable papers _"quite compelling."_ Nobody is claiming the peer-review replacement yet.
## Takeaways
1. **The paper is becoming an API.** Paper2Agent's bet: knowledge should ship as runnable, validated tools — not static text. 74 papers, 593 tools, 99% validation pass.
2. **Structure beats access.** Same model, same repos: 91.2% as callable tools vs 82.7% as files. The packaging of knowledge is now a first-class performance lever.
3. **Validation is the product.** The 3%-tolerance reproduce-the-paper gate + tool locking is what makes this trustworthy. Wrapping code in JSON-RPC is trivial; _proving the wrapper reproduces the science_ is the contribution.
4. **Agents compose.** Three paper-agents found GPR137 at rs887314 — a discovery none of the three papers made alone. The network of paper-agents is the real endgame.
5. **Agentification is a reproducibility certificate.** If your paper can't survive the pipeline, that's data about your paper.
_Sources: Nature (Sep 16, 2026) — "Reimagining research papers as interactive and reliable AI agents" (Miao, Davis, Zhang, Pritchard, Zou); AI Weekly coverage (Sep 24, 2026); IEEE Spectrum; Stanford Medicine. Numbers as reported by the authors._
**Companion notebook:** the runnable tutorial for this post — download it here (open in Colab/Jupyter).