Self-Driving Laboratories

Published

October 7, 2026

Self-driving labs, also called autonomous laboratories, close the experimental loop. A model proposes an experiment, automated hardware runs it, the result returns to the model, and the model proposes the next experiment. The autonomy is not in the robotics alone; it is in the loop. Evidence spans closed-loop chemistry and materials systems, while current biology examples remain lab-in-the-loop rather than autonomous robotic loops. Coscientist demonstrated the pattern for Pd-catalysed cross-coupling chemistry with GPT-4 as planner (Boiko et al., 2023). Virtual Lab demonstrated multi-agent biology research that produced experimentally validated SARS-CoV-2 nanobodies (Swanson et al., 2025). Robin extended the biology branch into lab-in-the-loop therapeutic-candidate discovery for dry age-related macular degeneration (Ghareeb et al., 2026). A-Lab reported autonomous discovery of dozens of inorganic materials and then drew published commentary and a later PRX Energy critique that illustrate how easily novelty claims can run ahead of validation (Szymanski et al., 2023; Neilson, 2023; Leeman et al., 2024). The capability is real for bounded optimisation; general autonomous discovery is not.

Nobel Turing Challenge

The Nobel Turing Challenge (Kitano, 2021) sets a different bar: an AI system that makes a discovery on a par with Nobel-level human science, fully or highly autonomously, with a commonly cited horizon around 2050. That frame separates the 2024 Nobels awarded for building AI systems (AlphaFold chemistry; neural-network physics) from a future prize for a discovery made by AI. A 2025 Nature News Feature surveys that distinction and records open expert timelines, from roughly a decade to fifty years, with some voices naming materials science or neurodegenerative disease as plausible first domains (Ahart, 2025). Challenge vision and expert opinion, not evidence that any deployed system has met the Nobel Turing bar or that open-ended autonomous biology is demonstrated.

Learning Objectives

Use this chapter to:

  • Distinguish laboratory automation, lab-in-the-loop research, and a closed autonomous experimental loop.
  • Evaluate a self-driving-laboratory claim using loop completeness, endpoint validity, exception handling, provenance, safety, and independent validation.

Prerequisites: Robotic Lab Automation and Cloud Labs for the hardware layer; Evaluation Principles for Life Sciences AI for the prospective-validation discipline.

Summary: Self-driving laboratories close the loop between hypothesis generation, experiment planning, robotic execution, measurement, and model update. Self-driving systems are real in selected chemistry and materials tasks; broad autonomous biology remains much more constrained.

Key point: Autonomy is bounded by protocol validity, instrument calibration, search space, safety rules, and whether the measured endpoint is meaningful. Open question: whether closed loops reproduce across laboratories, endpoints, organisms, and failure conditions.

Bottom line: Self-driving laboratories connect experimental design, automation, chemistry, synthetic biology, biomanufacturing, agents, and research governance.

Field Guide

What is this field trying to solve? Closing the loop between hypothesis generation, experiment planning, robotic execution, measurement, and model update.

What is the core idea? Autonomy is bounded by protocol validity, instrument calibration, search space, safety rules, and whether the measured endpoint is meaningful.

What is the current state of the field? Self-driving systems are real in selected chemistry and materials tasks; broad autonomous biology remains much more constrained.

What do we know, and what remains open? The evidence base includes Adam, Eve, Coscientist, Virtual Lab, Robin, A-Lab, Ada, ChemOS, mobile robotic chemists, cloud labs, optimization benchmarks, and automated assay platforms. The open question is whether closed loops reproduce across laboratories, endpoints, organisms, and failure conditions.

Why does this matter? Self-driving laboratories connect experimental design, automation, chemistry, synthetic biology, biomanufacturing, agents, and research governance.


Introduction

A self-driving laboratory has three parts: an inference loop that proposes the next experiment, hardware that executes the experiment, and an evaluation step that returns the result to the inference loop. The composition is older than current AI: Adam, built by Ross King and colleagues at Aberystwyth, demonstrated autonomous functional-genomics hypothesis testing in 2009 with logic-programming-based reasoning (King et al., 2009). Adam’s successor Eve focused on drug screening. The historical lineage matters because the conceptual contribution of self-driving labs predates the current LLM wave by more than a decade.

The current wave has three branches:

  • Chemistry and materials automation has peer-reviewed demonstrations for bounded tasks: thin-film optimization (Ada, MacLeod et al., 2020), mobile robotic chemistry (Burger et al., 2020), Pd-catalysed coupling (Coscientist, Boiko et al., 2023), and workflow orchestration (ChemOS, Roch et al., 2020). These studies do not establish production readiness across laboratories or reaction classes.
  • Materials autonomous discovery is a prominent example, with caveats: A-Lab’s reported 41 new inorganic materials in 17 days (Szymanski et al., 2023) drew published commentary and a later PRX Energy critique arguing that many “new” materials required stronger phase-identification and novelty evidence (Neilson, 2023; Leeman et al., 2024). The episode is a clear case study of how autonomous-discovery novelty claims can outrun validation.
  • Agentic biology research has peer-reviewed, lab-in-the-loop examples: Virtual Lab’s SARS-CoV-2 nanobody design with human-executed validation (Swanson et al., 2025) and Robin’s dry age-related macular degeneration candidate work (Ghareeb et al., 2026). These are adjacent to, but not evidence of, autonomous robotic biology loops.

Each branch is assessed below against the evidence framework, with the historical context that current claims are read against.

Demonstrated capability

Classify the system before judging the claim

The term self-driving laboratory should be reserved for a system that closes the inference, execution, measurement, and update loop. Robotics without model-directed iteration is laboratory automation. Agentic analysis with human-executed experiments is lab-in-the-loop research. Both can be valuable, but they support different claims.

Decision tree distinguishing laboratory automation, lab-in-the-loop research, and a closed self-driving laboratory.

Treat cloud access as orthogonal to autonomy: robotic execution of a fixed protocol is automation; model-directed next-step choice is autonomy; remote facility access is cloud. Batalis’s wet-lab autonomy vocabulary note (30 Sep 2026) is a short teaching companion beside this chapter’s automation / lab-in-the-loop / closed-loop distinction. Essay, not a closed-loop methods primary.

Before adoption or a public autonomy claim, record the following:

Record field Minimum evidence
Objective and endpoint Prespecified quantity the loop is optimizing and why it is scientifically meaningful
Search space Variables the system may change, fixed constraints, and out-of-scope actions
Physical execution Hardware, protocol version, calibration state, and evidence that the commanded procedure ran
Return path Measurement, quality-control rule, and how the result enters the next selection step
Human intervention Every manual correction, restart, interpretation, and exception-resolution step
Failure handling Invalid runs, missing measurements, instrument errors, abstention, and safe-stop behavior
Validation Independent confirmation of function or novelty, including negative and failed results
Reproducibility Complete loop logs, software and model versions, and transfer testing outside the originating setup

Use the downloadable self-driving laboratory loop log to separate platform throughput from validated discovery.

Chemistry automation: the mature branch

Coscientist is a peer-reviewed example of LLM-planned chemistry (Boiko et al., 2023). The system used language-model planning, web and documentation search, code execution, and automated experimental tools. The experimental case included Pd-catalysed cross-coupling reactions. The evidence supports orchestration of bounded, known chemical procedures, not general autonomous chemical discovery.

A practical self-driving lab is a design-make-test-analyse system, not only a model attached to instruments. Reviews of autonomous chemical experimentation emphasise that chemical production, characterization, calibration, and exception handling remain bottlenecks even when optimisation logic is automated (Seifrid et al., 2022). This is why the strongest papers report the physical loop, the control policy, and the validation result together.

A 2026 Nature Reviews Chemistry assessment identifies scalability, generalizability, and provenance-complete experimentation as the interdependent requirements for moving self-driving laboratories from focused demonstrations toward shared infrastructure (Canty and Abolhasani, 2026).

AutoLabs provides a current test of the agent-to-hardware boundary. The 2026 Scientific Reports study evaluated 20 agent configurations across five benchmark experiments and reported an F1 score above 0.89 on its challenging multi-plate chemistry benchmark when multi-agent orchestration and iterative self-correction were combined (Panapitiya et al., 2026). The system generated hardware-ready XML for one high-throughput liquid-handler platform. This is evidence for protocol translation and platform-specific execution, not autonomous discovery or transfer to other laboratory systems.

On 27 August 2026 Anthropic opened a research preview of the Model Hardware Standard, a model-agnostic driver (read/write primitives, device tags, MCP / CLI / API) so agents can operate programmable instruments (Anthropic, 2026). Partner notes include a Genentech BCA across liquid handler, arm, and plate reader, and a Baker/Pinglay agent-supervised qPCR. The spec is not open source; they state Claude still lacks physical intuition (foam vs software bug) and that safety evaluations are still being built. That is an instrument-interface preview, not a peer-reviewed self-driving lab. Instrument-access and biosecurity governance for autonomous agents is covered in the Biosecurity Handbook autonomous-agents chapter.

Ada (MacLeod et al., 2020) is an earlier and more focused example: a self-driving laboratory for thin-film materials discovery using Bayesian optimisation as the loop driver. The Science Advances paper reported optimisation of organic photovoltaic films with hundreds of automated cycles. The conceptual contribution was Bayesian-optimisation-driven autonomy at production scale before the LLM era.

ChemOS (Roch et al., 2020) is the orchestration layer beneath several of these systems. The PLOS ONE paper described an open-source workflow framework that handles experiment specification, scheduling, hardware integration, and result handling. ChemOS is a piece of infrastructure rather than a research result, and it appears in the methods sections of many subsequent self-driving lab papers.

The mobile robotic chemist (Burger et al., 2020) addressed a different constraint: untethered robotics that can move between instruments and conduct multi-instrument workflows. The Nature paper demonstrated an autonomous chemistry workflow that ran continuously for over a week and explored a photocatalyst search space.

Adjacent lab-in-the-loop agentic biology

Virtual Lab (Swanson et al., 2025) is a peer-reviewed example of multi-agent biology research. The Nature paper described a PI agent coordinating specialist agents to design 92 nanobodies against SARS-CoV-2 variants; selected designs were tested experimentally. Virtual Lab did not autonomously execute a robotic wet-lab feedback loop, so it should not be counted as self-driving-laboratory evidence.

Robin (Ghareeb et al., 2026) is the next major evidence point. It combined literature-search agents and data-analysis agents to generate hypotheses, propose experiments, interpret results, and update hypotheses. In the reported dry age-related macular degeneration workflow, Robin identified ripasudil and KL001 as candidates, proposed follow-up RNA-seq analysis, and generated the report’s main-text hypotheses, analyses, and figures. The important qualifier is the phrase lab-in-the-loop: humans still executed the wet-lab work, and replication remains the threshold for stronger claims.

Outside these systems, biology automation has been demonstrated for narrower tasks such as directed-evolution loops and assay optimization. Published autonomous-laboratory evidence is more extensive in chemistry and materials than in open-ended biology. The assay, hardware, metadata, and exception-handling layers all constrain transfer.

A-Lab and the novelty-claim discipline

A-Lab (Szymanski et al., 2023) reported 41 new inorganic materials autonomously produced and characterised in 17 days. The Nature 2023 paper was the highest-profile autonomous-discovery result of the year and drew immediate attention. Subsequent commentary and a later PRX Energy critique from solid-state chemists argued that novelty claims required stronger treatment of known phases, disorder, polymorphism, and automated X-ray diffraction interpretation (Neilson, 2023; Leeman et al., 2024).

The episode is the clearest current case study of the novelty-claim discipline that self-driving labs require:

  • Autonomous platform throughput is not autonomous discovery. Running 1,000 experiments per day produces 1,000 experiments per day. Whether any of them is a discovery depends on what was already known.
  • Automated characterization can be wrong in characteristic ways. Powder X-ray diffraction pipelines, machine-learning structure-solving, and database matching each have failure modes that look like new materials.
  • Independent expert review is part of validation. A novelty claim that does not pass scrutiny from the relevant subdiscipline is not a discovery.

A-Lab does not invalidate autonomous-laboratory work. It shows what the validation bar looks like.

Historical lineage: Adam, Eve, and the pre-LLM era

Adam (King et al., 2009), built at Aberystwyth, was the first robot scientist to autonomously generate functional-genomics hypotheses, design experiments to test them, and update its beliefs from results. The reasoning engine used logic programming and abductive inference. Adam ran for months with minimal human intervention and produced hypotheses about gene function in yeast that were independently verified. Eve, Adam’s successor, applied the same closed-loop architecture to drug screening.

The Adam and Eve work is conceptually upstream of the current LLM-driven systems. The closed loop, the integration with hardware, the question of what counts as discovery, and the validation discipline were all explored at Aberystwyth more than a decade before Coscientist. Current systems should be read in continuity with this lineage, not as a discontinuous emergence from LLM scaling.

Agentic chemistry orchestration: ChemCrow

ChemCrow (M. Bran et al., 2024) wraps an LLM around chemistry-specific tools (RDKit, retro-route planners, property predictors, web search, hardware interfaces). The Nature Machine Intelligence paper demonstrated tool-using agents on chemical synthesis planning tasks, drug-discovery workflows, and materials-design queries. ChemCrow is conceptually adjacent to Coscientist: where Coscientist is task-specific and hardware-coupled, ChemCrow is tool-coupled and more general. Together they describe two complementary approaches to LLM-driven chemistry workflows. The agentic-orchestration layer is covered in detail in Agentic Science Workflows.

GOLLuM fine-tunes a language encoder through a Gaussian-process objective so next-experiment selection carries calibrated uncertainty (Ranković et al., 2026). The Nature Machine Intelligence evaluation covers 23 published chemistry, materials, process, and molecular-design tasks, starting from ten low-performing conditions. GOLLuM is a calibrated optimizer over recorded outcomes, not a wet-lab self-driving laboratory.

Evidence anchor summary

Evidence Anchor What It Supports Practical Constraint
Coscientist LLM-planned autonomous Pd cross-coupling chemistry Reproduces known chemistry; novel-reaction discovery is separate
AutoLabs Agent-generated, hardware-ready protocols on a bounded chemistry benchmark One liquid-handler platform; protocol execution is not autonomous discovery
Virtual Lab Adjacent multi-agent biology research with human-executed experimental validation Not a robotic closed loop; one nanobody-design context
Robin Adjacent lab-in-the-loop agentic research Humans executed wet-lab work; replication and transfer remain open
A-Lab Autonomous inorganic materials platform at scale Novelty claims drew published commentary and PRX Energy critique; validation discipline required
Ada Bayesian-optimization-driven thin-film discovery Pre-LLM workflow at production scale
ChemOS Open-source orchestration layer for self-driving chemistry Infrastructure, not a research result by itself
Burger mobile chemist Untethered robotic chemistry across instruments Single laboratory, single chemistry domain
ChemCrow LLM-orchestrated chemistry tools Tool-coupled rather than hardware-coupled
GOLLuM Uncertainty-calibrated LLM optimizer for next-experiment selection Benchmark optimizer over recorded outcomes; not a robotic wet-lab loop
Adam (King 2009) First closed-loop robot scientist (functional genomics) Logic-programming reasoning, not deep-learning
Self-driving lab review (Canty 2026) Current field assessment Scalability, generalizability, and provenance remain infrastructure constraints (Canty and Abolhasani, 2026)

Theoretical capability

Several capabilities are plausible but not yet routine.

Closed-loop biology across targets and assays. Biology assays can be slow, variable, and difficult to automate. Virtual Lab and Robin are agentic, lab-in-the-loop examples rather than autonomous robotic loops. Prospective, repeated closed-loop biology across multiple targets and assay classes remains an open evidence need.

A 2026 preprint on closing the loop in biomedical discovery argues that progress is constrained as much by the verification budget as by hypothesis quality, and that hard wet-lab verifiers are not interchangeable with soft simulators (Fang et al., 2026). Fuller teaching is on Agentic Science Workflows. Preprint perspective, not evidence that any deployed agent closes the loop.

Autonomous discovery of new reaction mechanisms. Reproducing known chemistry at high throughput differs from discovering new chemistry. Current systems can explore parameter spaces inside a defined reaction class; moving to a new reaction class typically requires human chemists to define the search space.

Cross-laboratory reproducibility. Independent reproduction on a materially different autonomous platform would provide stronger evidence of transfer. Current work remains dominated by single-platform and single-laboratory demonstrations.

Autonomous in-vivo experimentation. Cell-free and in-vitro experiments dominate current self-driving lab work. In-vivo experimentation adds welfare, protocol, review, monitoring, assay, and robotics constraints. These are governance and technical limits, not one interchangeable bottleneck.

Autonomous goal-setting. Current systems take a human-defined goal (optimize this property, find a binder to this target) and execute toward it. Setting the goal itself, deciding which problem to work on, remains human. Whether this is a desirable target for autonomy is an open question.

On the path from question-answering tools to question-posing agents, Guo et al. place autonomous laboratories inside a broader agentic-science trajectory and name hallucination, bias, dual use, and cognitive deskilling as core risks, arguing that progress still depends on responsible human-AI collaboration (Guo et al., 2026). Perspective on trajectory and governance, not a closed-loop SDL validation study.

Beyond current capability

A few framing claims are not supported by current evidence.

Self-driving labs make scientists obsolete. They do not. Goal definition, experimental judgment, novelty assessment, and integration across subdisciplines remain human. The autonomous platforms accelerate the parts of science that are well-characterized; they do not replace the parts that are not.

Autonomous labs reliably discover new biology without expert oversight. Current evidence does not support that claim. The A-Lab episode illustrates that high-throughput platforms can produce candidate discoveries that require expert review and independent validation. Validation is part of the scientific claim, not a downstream formality.

Closed-loop optimization generalizes across science. It does for narrow, well-characterized optimization problems (thin films, photocatalysts, reaction yields). It does not yet for open-ended discovery, where the right question is part of the problem.

Robot scientists are a recent invention. They are not. Adam and Eve at Aberystwyth demonstrated the closed-loop concept in 2009. Current LLM-driven systems are an architectural advance over Adam’s logic programming, not a discontinuous new capability.

Evidence that would change the assessment

Autonomous-lab claims become more promising when they reproduce across independent facilities, with prespecified objectives, complete loop logs, and external confirmation of novelty or function. For biology, the threshold is prospective closed-loop optimization that improves a validated assay across multiple targets or cell systems without hiding exception handling, failed loops, or human intervention.

Implications for research and program decisions

For researchers and program leaders evaluating self-driving lab claims:

  • Count the loop iterations. A platform that ran once is automation; a platform that ran a closed inference-experiment-update loop many times is autonomy. The number of loop iterations should be in the paper.
  • Read the validation layer. Identify which experiments verified novelty, correctness, and reproducibility. A-Lab’s experience shows what happens when this layer is thin.
  • Match the chemistry-to-biology gap. Bounded chemistry and materials loops have more mature demonstrations than autonomous biology loops. A platform that works for thin films does not directly transfer to nanobody discovery, and vice versa.
  • Read the hardware layer alongside the algorithm layer. Cloud-lab integration (Emerald, Strateos), robotic-arm reliability, and assay-format constraints often dominate what is possible. The published algorithm sits on top of a hardware stack that bounds its real performance.
  • Treat single-laboratory results as proof of concept, not as a community capability. A self-driving result that works in one lab and does not reproduce in another has demonstrated something narrower than the abstract suggests.
  • Cite the historical lineage. Current self-driving lab papers that ignore Adam, Eve, Ada, ChemOS, and the Häse 2019 review tend also to overclaim novelty. The honest framing locates current work in the trajectory.
  • Pair this chapter with Agentic Science Workflows. The orchestration-agent layer (ChemCrow, Coscientist’s planner, Virtual Lab’s PI-agent) is conceptually distinct from the autonomous-loop layer, even though both appear in the same papers.