An artificial intelligence system called Biomni completed biomedical research tasks with 60% accuracy, matching the performance range of human experts (60% to 70%), in roughly three minutes per analysis, according to a paper published in Science on July 9.
The study, led by Kexin Huang at Stanford University, evaluated Biomni across benchmarks spanning genetics, genomics, and microbiology. On the primary evaluation reported in Science, the system achieved 60% accuracy across cases, completing analyses in approximately three minutes compared with hours for human researchers. In a separate evaluation detailed in a preprint hosted by PubMed Central, Biomni achieved 74.4% accuracy against expert human performance of 74.7%, outperforming all baseline models.
The paper is the first demonstration of a single AI agent executing the full biomedical research workflow—literature search, hypothesis generation, experimental design, and data analysis—within a unified system. Prior investigations into autonomous AI research systems have tracked the technology from specialized mathematical agents to general-purpose science tools, with systems like Aletheia demonstrating autonomous research paper generation and open problem-solving in mathematics.
However, the reliability of AI-generated analysis remains an open question. A study published in PNAS found that AI agents analyzing identical datasets produced wildly different conclusions depending on their behavioral prompts raising concerns about whether autonomous agents can be trusted for definitive research, even when their methods are technically correct. The editor of Science has warned that AI is making scientific publishing slower, costlier, and less reliable, as submission volumes have risen 42% since ChatGPT's release while writing quality has measurably declined.
Data on how Biomni performs on tasks outside its curated benchmarks was not reported. The paper does not address clinical deployment, liability, or regulatory pathways.