Accuracy on three kinds of query, comparing a converted paper against a coding agent that already has the repository, and against Biomni. Higher is better. Hover any bar for the exact figure.
| Query type | Paper2Agent | Claude + Repo | Biomni |
|---|---|---|---|
| 15 tutorial-derived queries | 98.7% | 82.7% | 37.3% |
| 15 novel queries | 100.0% | 78.7% | 56.0% |
| 30 open-ended queries | 82.7% | 56.7% | 72.2% |
The pattern worth reading: the gap is narrowest on open-ended questions, where Biomni closes in. Structured, procedural questions — the ones a paper documents best — are where a validated tool set pulls furthest ahead.
Figures as reported for Paper2Agent (Nature, September 16, 2026) by MarkTechPost, from the authors' own benchmarks. Treat vendor and author benchmarks as directional, not definitive.