Interactive · Benchmark results

Paper2Agent answers paper questions accurately — and widens its lead on new ones

Accuracy on three kinds of query, comparing a converted paper against a coding agent that already has the repository, and against Biomni. Higher is better. Hover any bar for the exact figure.

Paper2Agent Claude + Repo Biomni
Accuracy, % (mean across five runs; graded by two human experts at 96.7% inter-rater agreement)
Query typePaper2AgentClaude + RepoBiomni
15 tutorial-derived queries98.7%82.7%37.3%
15 novel queries100.0%78.7%56.0%
30 open-ended queries82.7%56.7%72.2%

The pattern worth reading: the gap is narrowest on open-ended questions, where Biomni closes in. Structured, procedural questions — the ones a paper documents best — are where a validated tool set pulls furthest ahead.

Figures as reported for Paper2Agent (Nature, September 16, 2026) by MarkTechPost, from the authors' own benchmarks. Treat vendor and author benchmarks as directional, not definitive.