AI 'Co-Scientist' Matches Human Experts in Minutes

AI 'Co-Scientist' Matches Human Experts in Minutes

Stanford's Biomni automates the full biomedical research workflow—literature search, hypothesis, experiment design—matching expert accuracy in 3 minutes vs. hours.

LS
Linsey Smith
Aug 5, 2026
1 min read

An artificial intelligence system called Biomni completed biomedical research tasks with 60% accuracy, matching the performance range of human experts (60% to 70%), in roughly three minutes per analysis, according to a paper published in Science on July 9.

The study, led by Kexin Huang at Stanford University, evaluated Biomni across benchmarks spanning genetics, genomics, and microbiology. On the primary evaluation reported in Science, the system achieved 60% accuracy across cases, completing analyses in approximately three minutes compared with hours for human researchers. In a separate evaluation detailed in a preprint hosted by PubMed Central, Biomni achieved 74.4% accuracy against expert human performance of 74.7%, outperforming all baseline models.

The paper is the first demonstration of a single AI agent executing the full biomedical research workflow—literature search, hypothesis generation, experimental design, and data analysis—within a unified system. Prior investigations into autonomous AI research systems have tracked the technology from specialized mathematical agents to general-purpose science tools, with systems like Aletheia demonstrating autonomous research paper generation and open problem-solving in mathematics.

However, the reliability of AI-generated analysis remains an open question. A study published in PNAS found that AI agents analyzing identical datasets produced wildly different conclusions depending on their behavioral prompts raising concerns about whether autonomous agents can be trusted for definitive research, even when their methods are technically correct. The editor of Science has warned that AI is making scientific publishing slower, costlier, and less reliable, as submission volumes have risen 42% since ChatGPT's release while writing quality has measurably declined.

Data on how Biomni performs on tasks outside its curated benchmarks was not reported. The paper does not address clinical deployment, liability, or regulatory pathways.

About the Writer

More from Mindplex

Keep reading

Three more ideas worth your time.

Browse MindBytes

Discussion

Join the discussion

Sign in to share a response with the community.

Type @ to mention someone Type / or use + to add a block Highlight text, then choose Link
Loading editor

Comments cannot be edited after posting because they become part of the reputation record. Give yours a quick review first.