Agentic Science Workflows
Agentic science workflows use AI agents to plan, retrieve, write, use tools, and coordinate tasks. In life sciences, the central issue is not whether an agent sounds scientific. The issue is whether it preserves provenance and respects experimental limits. ChemCrow, Coscientist, Virtual Lab, and CellVoyager are peer-reviewed examples across chemistry, wet-lab biology coordination, and computational biology; ARPA-H IGoR places agentic systems inside a governed research-infrastructure program. The discipline is bounded tasks, source control, audit logging, separated permissions, and explicit human authorization gates for actions involving biological materials.
Use this chapter to:
- Use tool-calling and multi-agent systems to plan, retrieve, analyze, and coordinate parts of scientific work without losing provenance or human control.
- The tool boundary matters: literature retrieval, code execution, database queries, laboratory actions, and biological design each need different gates.
Prerequisites: Self-Driving Laboratories for the hardware-coupled case; Information Hazards in Capability Research for the dual-use dimension of autonomous biological actions.
Summary: Use tool-calling and multi-agent systems to plan, retrieve, analyze, and coordinate parts of scientific work without losing provenance or human control. Agents can help with bounded research workflows, but fully reliable autonomous science remains limited by tool validation and biological evidence.
Key point: The tool boundary matters: literature retrieval, code execution, database queries, laboratory actions, and biological design each need different gates. Open question: whether agents can complete auditable workflows while preserving human authorization for consequential steps.
Bottom line: Agentic workflows connect literature, datasets, computation, lab automation, governance, information hazards, and program operations.
What is this field trying to solve? Use tool-calling and multi-agent systems to plan, retrieve, analyze, and coordinate parts of scientific work without losing provenance or human control.
What is the core idea? The tool boundary matters: literature retrieval, code execution, database queries, laboratory actions, and biological design each need different gates.
What is the current state of the field? Agents can help with bounded research workflows, but fully reliable autonomous science remains limited by tool validation and biological evidence.
What do we know, and what remains open? Known reference points include ChemCrow, Boiko Coscientist, Google Co-Scientist, FutureHouse Robin, Virtual Lab, Virtual Biotech, CellVoyager, ARPA-H IGoR, tool-use benchmarks, provenance logs, and workflow orchestration systems. What remains open is whether agents can complete auditable workflows while preserving human authorization for consequential steps.
Why does this matter? Agentic workflows connect literature, datasets, computation, lab automation, governance, information hazards, and program operations.
Introduction
ARPA-H IGoR names AI/ML orchestration and agentic systems alongside laboratory automation, protocol standardisation, and distributed systems (ARPA-H IGoR, 2026). That framing puts agents inside a governed research infrastructure rather than outside it. The canonical published systems are ChemCrow (M. Bran et al., 2024), Boiko Coscientist for chemistry (Boiko et al., 2023), Virtual Lab (Swanson et al., 2025), Google Co-Scientist for biomedical hypothesis generation (Gottweis et al., 2026), FutureHouse Robin for lab-in-the-loop candidate discovery (Ghareeb et al., 2026), and CellVoyager for autonomous single-cell analysis (Alber et al., 2026).
The governance vocabulary already exists. NIST’s AI Risk Management Framework separates governance, mapping, measurement, and management of AI risk (NIST, 2023). WHO’s health AI guidance stresses human oversight, transparency, accountability, and equity for health-related AI systems (WHO, 2021). In life-sciences research, those ideas translate into scoped permissions, logged tool calls, source-linked claims, and human authorisation before any action that touches biological materials.
Agent architecture for research work
A research agent should be described by its tools, permissions, memory, retrieval corpus, execution environment, and audit record. A literature-only agent needs source retrieval, claim extraction, citation verification, and reviewer signoff. A computational-analysis agent needs sandboxed code execution, data-version control, notebook replay, and package pinning. A lab-orchestration agent needs protocol boundaries, hardware interfaces, reagent ordering limits, and human authorization gates.
The same model may sit inside each architecture, but the risk changes with the tool boundary. Read access creates source-quality risk. Code execution creates reproducibility and security risk. Procurement creates cost and safety risk. Laboratory execution creates biological and institutional risk. Evaluation should therefore be permission-specific rather than agent-generic.
A 2026 Nature Biotechnology perspective maps the emerging architecture of biomedical agent teams and the deployment problems that follow from autonomy, memory, tool use, and inter-agent coordination (Li et al., 2026). It is a field framework, not comparative evidence that agent teams improve scientific outcomes. Agentic architecture and demonstrated research utility remain separate claims.
Tool validation and provenance records
Agentic systems should produce a provenance record that a reviewer can replay: prompt or task, retrieved sources, tool calls, parameters, input data versions, intermediate files, generated code, execution logs, error messages, human approvals, and final outputs. A polished report without the record is weak evidence because the reasoning path cannot be checked.
Tool validation is separate from model evaluation. A PubMed tool should be checked for recall and citation alignment. A chemistry tool should be checked against known inputs and invalid outputs. A notebook tool should be checked for deterministic replay. A cloud-lab tool should be checked against protocol simulation and physical run logs. Agents are only as reliable as the least-audited tool in the chain.
Configured-System Evaluation and Repeated Runs
The evaluated object is the configured agent system, not the model name alone. Model or API version, system prompt, tools, retrieval and memory, permissions, environment state, and attempt, time, and cost budgets can all change the result. NIST’s initial public draft on automated benchmark evaluations treats model and agent settings as part of the evaluation protocol and calls for explicit objectives, reproducible runs, uncertainty analysis, and qualified reporting (NIST AI 800-2, initial public draft, 2026).
Agent evaluation should combine outcome, process, and safety evidence. Outcome checks ask whether the intended scientific state was reached. Process checks examine tool selection, provenance, recovery from errors, and unnecessary repetition. Safety checks examine unauthorized actions, disclosure, permission escalation, and whether required human gates fired. In the non-biological customer-service domains of tau-bench, final database state and repeated trials exposed failures that single conversational scores missed (Yao et al., 2025). The evaluation design is informative for life-sciences systems, but the benchmark results do not establish biological performance.
Record the following for each release candidate:
- System identity: model or API date, system prompt, scaffold, tool versions, retrieval corpus, memory policy, and environment image
- Authority: data access, code execution, procurement, communication, and laboratory permissions, including every human approval gate
- Evaluation conditions: task set and split, initial state, attempt and time budget, sampling settings, and comparator
- Measures: scientific outcome, provenance completeness, policy compliance, side effects, recovery behavior, human intervention, latency, and cost
- Reliability: repeated-run distribution and uncertainty, not only the best or average attempt
- Evaluator validity: deterministic checks where available and blinded domain-expert calibration for model- or rubric-based judgments
- Failure custody: retained traces, severity, owner, regression-test identifier, and resolution status
- Reassessment triggers: any change to the model, prompt, tools, retrieval, memory, permissions, environment, or evaluator
Model judges are useful when deterministic verification is unavailable, but they require their own validation. Human and LLM judges both showed susceptibility to multiple forms of judgment bias in an EMNLP study (Chen et al., 2024). Calibrate a judge against blinded domain experts, retain disagreements, and keep the final gate independent of the model family being evaluated where feasible.
Every confirmed failure should enter a held-out regression suite. Re-run that suite after any change to the configured system, and use multiple trials for stochastic workflows. Do not train the policy or evaluator on the final regression set, because a memorized fix is not evidence of general reliability.
Human gates for biological actions
The human gate should sit before any action that orders materials, reserves lab time, modifies a biological protocol, executes an experiment, or shares dual-use relevant details. The gate should approve a specific action, not a broad agent role. “This agent may read PubMed” and “this agent may order a construct” are different permissions and should expire separately.
Demonstrated capability
Demonstrated capability includes literature triage, code execution, protocol drafting, bounded workflow orchestration, semi-autonomous biomedical hypothesis generation, lab-in-the-loop candidate discovery, and autonomous single-cell analysis. ChemCrow, Boiko Coscientist, Virtual Lab, Google Co-Scientist, FutureHouse Robin, CellVoyager, and IGoR provide the current evidence anchors (M. Bran et al., 2024; Boiko et al., 2023; Swanson et al., 2025; Gottweis et al., 2026; Ghareeb et al., 2026; Alber et al., 2026; ARPA-H IGoR, 2026).
For a peer-reviewed gene-editing co-pilot with wet-lab knockout and CRISPRa showcases, see CRISPR-GPT on Cell and Gene Therapy (Qu et al., 2025).
| Evidence Anchor | What It Supports | Practical Constraint |
|---|---|---|
| ChemCrow | LLM-tool orchestration for chemistry | Tool selection and validity bound output |
| Coscientist | LLM-planned closed-loop chemistry | Bounded chemistry domain; reproduces known chemistry |
| Virtual Lab | Multi-agent biology research with wet-lab validation | One target; generalisation is open |
| Google Co-Scientist | Biomedical hypothesis generation with in vitro validation | Hypotheses still require expert selection and experimental testing |
| FutureHouse Robin | Lab-in-the-loop candidate discovery and follow-up data analysis | Human execution, biological interpretation, and replication remain necessary |
| CellVoyager | Autonomous computational analysis of scRNA-seq datasets | Notebook analyses still require expert review and reproducibility checks |
| ARPA-H IGoR | Agentic systems inside governed infrastructure | Program ambition is not proof of deployed reliability |
| Bridge2AI | AI-ready data and workforce materials | Agents need high-quality inputs and human review |
| EMA and FDA materials | Lifecycle accountability for AI in regulated contexts | Regulatory use requires documentation |
| NIST AI RMF | Govern-map-measure-manage risk structure | General framework; must be adapted to biology |
| WHO health AI guidance | Human oversight, transparency, accountability, equity | Health framing does not replace lab-specific biosafety review |
In September 2026 Anthropic reported that Claude Mythos 5 agents, coordinated in a worker–supervisor orchestration from a high-level reverse-transcriptase genome-mining brief, surveyed on the order of 200,000 RT clusters across roughly 1.9 billion protein clusters and returned ranked human-readable reports after about 21.5 hours of wall-clock work (~949 sessions, ~216 million tokens) without mid-campaign human interruption (Anthropic news, 23 Sep 2026; technical preprint PDF). One agent reading raw flanking DNA flagged an unannotated tandem-repeat array beside a jumbo-phage RT; the resulting system, array-associated reverse transcriptases (ART), couples an RT, a dedicated partner gene, and evenly spaced DNA repeats expressed as distinct short RNAs. Anthropic states that function remains unknown and that wet-lab characterization in its Bay Area lab is done by human scientists under BSL-1/BSL-2 only. Company preprint and news; not a peer-reviewed article; not evidence that ART is a gene editor or that agents may authorize biological materials without a human gate.
Paper2Agent automatically converts a research paper plus its codebase into a Model Context Protocol (MCP) server of tools, resources, and prompts, then connects that server to a chat agent so natural-language queries invoke the paper’s methods rather than requiring readers to adapt the codebase by hand (Miao, Davis, Zhang, Pritchard, Zou, 2026; code). Case studies cover AlphaGenome, Scanpy, and TISSUE: agents reproduced tutorial results, answered novel queries, and collaborated across paper MCPs on genomic interpretation (including a psoriasis causal-gene prioritization demo). The published paper reports AlphaGenome MCP construction in around 45 minutes (US$14) on a personal laptop without human intervention, and Scanpy preprocessing tools in around 45 minutes (US$13). This is dissemination and reuse infrastructure for methods papers with usable code, not evidence that agent teams raise wet-lab discovery rates, and agentification fails when the codebase is incomplete.
Paper-as-tool MCP servers belong in the same operational inventory as other research workbenches; see Toolkit for AI-Augmented Bio Research.
Building on the Virtual Lab idea, Zhang, Zou, and colleagues describe a Virtual Biotech: a chief-scientific-officer agent coordinating domain scientist agents across genetics, omics, chemoinformatics, and clinical data (Zhang et al., 2026). In one showcase, more than 37,000 clinical-trialist agents curated outcomes from 55,984 trials and linked targets to single-cell features; drugs aimed at cell-type-specific, switch-like genes were associated with higher phase advancement and fewer adverse events in their analysis, and separate case studies proposed a B7-H3 antibody–drug conjugate strategy for lung cancer and failure modes for a terminated ulcerative colitis trial. Agent-curated observational associations and design case studies: not wet-lab proof that agent companies raise approval rates.
Theoretical capability
Theoretical capability includes agents that propose experiments, call analysis tools, update models, and prepare protocol-ready plans. This is plausible for bounded settings with source control, tool permissions, and human approval. Generalisation beyond bounded settings to open-ended scientific autonomy is a research frontier rather than a deployed capability.
The operational threshold is not whether an agent can draft a protocol. It is whether the protocol, data sources, tool versions, intermediate decisions, and human approvals are recoverable after the fact. Agentic workflows that cannot produce that record are not suitable for capability claims, publication decisions, or regulated work.
Guo and colleagues map biomedical AI’s move from tools that answer human-posed questions to agents that pose their own, assessing scaling logic and the autonomous-laboratory trajectory alongside risks of hallucination, bias, dual use, and cognitive deskilling (Guo, Ting, Su, Aliper, Zhavoronkov, and Ting, 2026). Their conclusion centers responsible human-AI collaboration, not unsupervised wet-lab agency. Field perspective on agentic science and named risks, not authorization for ungoverned laboratory agents or evidence that autonomy already replaces scientific judgment.
Beyond current capability
Beyond current capabilities includes unsupervised agents conducting open-ended biological research without human governance. Biological materials, safety controls, privacy, and scientific accountability require explicit human authority. A May 2026 Nature editorial on “AI scientists” framed the boundary correctly: process and speed do not replace the human judgment that makes research worth doing and keeps it accountable (Nature, 2026). AI scientists that operate without expert review remain beyond current capabilities and outside biosafety norms.
A 2026 preprint on closing the loop in biomedical discovery driven by AI argues that “AI scientist” progress is constrained as much by the verification budget as by hypothesis quality, and that evaluation should test experiment selection, revision under new evidence, and calibrated uncertainty across soft (simulation/surrogate) and hard (wet-lab) verifiers (Fang et al., 2026). Preprint perspective, not a peer-reviewed article, and not evidence that any deployed agent closes the loop.
A September 2026 Google and Google DeepMind early-insights snapshot triangulates about 15 million Gemini interactions, an inventory of over 2,600 specialized AI-for-science models, and a survey of 637 US and UK scientists (Codreanu, Imas, Mateos-Garcia et al., 2026). Scientists report high adoption (nearly half using some form of AI daily) and average time savings of about 6.9 hours per week, mostly reinvested in research, while LLMs and specialized models appear complementary rather than interchangeable. The same survey finds bottlenecks shifting downstream: about 41% report a larger backlog of untested hypotheses, and among those who save time about 46% spend more than a quarter of that saved time verifying AI outputs, with the verification burden described as particularly high in the life sciences (Codreanu et al., 2026). This is institutional observational evidence with self-report and Gemini-telemetry limits, not causal proof that AI has raised discovery rates or closed the experimental loop.
OpenAI’s September 2026 research-acceleration snapshot claims an “automated research intern” level inside its lab and rising agent-workday use under human priority-setting (OpenAI, 2026). Vendor self-report; see the Biosecurity Handbook for the fuller RSI and control-pause teaching.
DeepMind’s GDM AI Control Roadmap (v0.1) adds system-level detection and response ladders for imperfectly aligned internal agents (Phuong et al., 2026). Vendor roadmap; see the Biosecurity Handbook for the fuller control teaching.
Evidence that would change the assessment
Agentic workflows become more promising when agents complete bounded research workflows with replayable source records, verified tool calls, and independent recovery from errors across tasks. Stronger claims need prospective evaluation showing that agent outputs change experimental decisions while preserving human authorisation for procurement and laboratory execution.
Implications for research and program decisions
- Require source links for literature-derived claims.
- Log tool calls, parameters, data versions, and outputs.
- Freeze the configured system in a System Evaluation Card and evaluate outcome, process, safety, and repeated-run reliability before release.
- Convert confirmed failures into a held-out regression suite and rerun it after changes to the model, prompt, tools, retrieval, memory, permissions, environment, or evaluator.
- Use separate permissions for reading, analysis, procurement, and laboratory execution.
- Block autonomous actions involving biological materials unless a human authorises the exact protocol.
- Cite ChemCrow, Coscientist, and Virtual Lab specifically rather than referring to “AI agents” generically.
- Treat ARPA-H IGoR as programmatic infrastructure, not as a specific deployed system.