A general-purpose biomedical AI agent called Biomni was published today in the journal Science, with authors reporting that the system can decompose research problems, select and combine computational and laboratory tools, and complete a range of biomedical tasks with performance approaching human experts in some real-world studies.
What the team built
The project, led by Chinese scientist Huang Kexin and collaborators, describes two coupled elements: Biomni-E1, the execution environment that integrates commonly used scientific tools, databases and software; and Biomni-A1, the agent that operates within that environment. Unlike prior agents confined to narrow domains, the authors say Biomni was designed to work across multiple biomedical areas without relying on a single fixed workflow.
Capabilities and demonstrations
According to the paper, Biomni shows strong generalization across tasks in genetics, genomics and pharmacology. The team reports that in selected real scientific problems the agent's performance was close to that of human specialists and that it often completed analyses in less time. The paper also presents case studies in which Biomni:
- interpreted multimodal datasets;
- optimized protein stability;
- coordinated wet-lab instrument operations; and
- generated experimental protocols that could be experimentally validated.
“Biomni does not rely on a fixed workflow template,”
The authors emphasize that the agent can autonomously break down tasks and call tools as needed, rather than following a rigid, preprogrammed pipeline.
Data curation and scale
To define an action space and train discovery mechanisms, the team analyzed literature across 25 biomedical topics defined by bioRxiv. They sampled 100 recent papers from each topic, yielding a corpus of 2,500 papers that informed the agent's understanding of tasks, software and databases.
| Item | Count |
|---|---|
| Topics analyzed | 25 |
| Papers per topic | 100 |
| Total papers | 2,500 |
Implications and caution
The authors propose that agents like Biomni could augment human researchers, accelerating the translation of basic findings into applications by assisting with complex experimental planning and execution. They argue this represents “a promising direction” for collaborative human–AI research. At the same time, the paper documents technical validation rather than broad deployment; real-world use will require attention to reproducibility, safety, laboratory oversight, and validation against independent datasets and human judgment.
The full manuscript is available through Science (DOI: 10.1126/science.adz4351), where readers can assess detailed methods, benchmarks and case studies.
As biomedical AI systems grow more capable, institutions, funders and regulators in the United States and elsewhere will face decisions about validation standards, access to integrated toolchains, and how to integrate such agents into lab workflows while maintaining scientific rigor and patient safety.