Executive Summary

Published

August 22, 2026

Life sciences AI now reaches beyond isolated modelling tasks into biological discovery, translational research, and research operations. Protein structures are predicted before they are solved. Cells are represented by foundation models. Organismal and environmental systems are becoming model targets. Generative systems can propose more molecules than a laboratory can synthesize and test. Autonomous chemistry runs closed loops on robotic platforms. The practical risk is not only overclaiming. It is using the wrong evidence for the wrong decision.

Core Takeaways

The handbook organises evidence into three tiers throughout. The summary below names the bottom-line state of each major area.

Protein structure prediction is a routine research input. AlphaFold 2 produced a marked improvement on many single-chain targets at CASP14, with accuracy competitive with experimental structures for much of the benchmark. AlphaFold 3 extended prediction to biomolecular interactions involving nucleic acids, ions, and small-molecule ligands. DeepMind released AlphaFold 3 inference code and weights for academic, noncommercial use in November 2024, while Boltz and Chai provided separately developed models under different licenses. Structure prediction is now routine research infrastructure, but confidence scores and experimental validation determine which downstream claims are defensible.

Protein design is an experimental discipline supported by generative models. RFdiffusion generates backbones. ProteinMPNN and LigandMPNN design sequences. ESM3, Chroma, and ProGen2 add multimodal and autoregressive generative approaches. RFantibody (Nature 2026) extends the lineage to peer-reviewed de novo antibody design. AlphaProteo from DeepMind reports strong binder generation but remains an unpublished arXiv preprint with restricted code. Every designed protein still needs expression, purification, characterisation, functional assay, and developability review before it counts as a candidate.

Single-cell foundation models do part of what their papers claim. scGPT, Geneformer, scFoundation, and scBERT produce reusable cell-state representations. They support cell-atlas search and selected integration tasks. Nicheformer provides peer-reviewed evidence that spatial context can be incorporated during pretraining. Independent evaluations published in 2024 and 2025 show that on perturbation prediction and several core tasks, the deep approaches do not consistently outperform PCA plus linear regression. The evidence supports single-cell foundation models for some tasks and leaves others unverified.

Genome foundation models are early and active. Evo demonstrated sequence modelling across DNA, RNA, and proteins. Evo 2 extended genome modelling across all domains of life, with open model parameters, code, and OpenGenome2 data. Nucleotide Transformer, GET, Orthrus, and AlphaGenome show the field splitting into genome-wide, transcriptional, RNA-specific, and regulatory-variant models. Production-quality variant interpretation still happens inside ACMG/AMP and ClinGen frameworks, with computational predictors as evidence inputs rather than classifications.

Systems-biology claims require causal discipline. Cell-state embeddings, pathway scores, and network diagrams become useful when they name a mechanism and a falsifying perturbation. Gene regulatory network inference and whole-cell modeling are established research traditions, but observational edges are not causal maps. Virtual-organism claims are not larger virtual-cell claims; they require development, physiology, tissue coupling, environment, time, and prospective validation.

Organismal and environmental biology belong in the core scope. Neuroscience AI, aging clocks, plant and crop AI, ecological monitoring, environmental DNA, and virtual organisms are not side cases. They are the scale rung between cellular biology and translation. The evidence problem is measurement validity: whether the model is learning biology or platform artifacts, geography, sampling effort, season, field conditions, or cohort structure.

AI for therapeutics is delivering candidates, not proof of improved approval probability. Rentosertib reached a randomized Phase 2a trial whose primary endpoint was safety and tolerability; forced vital capacity was exploratory and the study was too small and short to establish registration-level benefit. Public evidence reviewed on August 12, 2026 did not identify an FDA- or EMA-approved small molecule whose molecular discovery was attributed to generative AI. That time-stamped negative claim must be rechecked as the pipeline changes. Comparative evidence that AI improves clinical success rates, attrition, or total program cost remains insufficient.

Translation is broader than small molecules. Chemical biology, target engagement, cell and gene therapy, diagnostics, biomarkers, trials, and real-world evidence each have different validation objects. A generated molecule, engineered cell, biomarker signature, or diagnostic feature becomes useful only when the evidence matches the decision it is supposed to support.

Self-driving laboratories and agentic systems work for bounded tasks, but they are not the same category. Self-driving laboratories close a machine-directed experimental loop through robotic execution. Coscientist and A-Lab provide chemistry and materials examples. Virtual Lab, Co-Scientist, and Robin are agentic or lab-in-the-loop research systems in which people perform or authorize key experimental steps. Open-ended autonomous biology beyond bounded settings remains beyond current capabilities.

Information hazards are handled through deliberate disclosure and current institutional review. The standard is enough detail for scientific verification without unnecessary operational detail that raises misuse risk. The July 2026 USG Policy for Stopping High-Risk Life Sciences Research replaced the prior transition state, prohibited federal support for defined dangerous gain-of-function research, established international-research restrictions, and required agency implementation guidance. NIH states that previously paused activities remain paused until NIH-specific requirements are established.

Biosecurity evaluation requires more than a capability benchmark. Scientific utility, hazardous capability, human or agent uplift, safeguard effectiveness, operational consequence, and lifecycle governance are different claims. A public benchmark can support one of them without establishing the others. Release decisions should therefore use configuration-specific testing, matched baselines, statistical uncertainty, legitimate-use evaluation, adversarial safeguard testing, independent review, and post-deployment monitoring. The Benchmarks for Bio AI chapter provides the evaluation framework.

Reproducibility is computational and experimental. DOME is a useful community checklist for reporting supervised machine learning in biology; it is not a universal professional standard. Model cards and dataset cards add documentation. Bridge2AI and ARPA-H IGoR provide federal infrastructure examples. A notebook that reruns is necessary but not sufficient if the assay cannot be repeated.

Workforce and compute are institutional, not individual. A credible team is cross-disciplinary: biological expertise, data engineering, ML, wet-lab partnership, regulatory engagement when relevant, and governance. Compute without experimental judgement creates expensive noise.

Reading Rule

Every major capability claim in the handbook lands in one of three tiers:

  • Demonstrated for claims supported by peer-reviewed evidence, official documentation, or reproducible blind benchmarks.
  • Theoretical for claims plausible under current methods but not yet established for routine use.
  • Beyond current capabilities for claims not supported by credible evidence with current systems.

If a claim cannot be placed in a tier, the claim is not specific enough.

Practical Use

The handbook is built to be read in the situations that matter:

  • Evaluating a model claim before adopting it in a research program.
  • Designing a validation plan that matches the biological decision the model informs.
  • Briefing a research team on the current evidence for a given model class.
  • Reviewing a vendor pitch against the published evidence in peer-reviewed venues.
  • Deciding whether a newly published paper changes a program decision.

It is not a tutorial in machine learning. It is a reference for working biologists, biotechnology leaders, physician-scientists, drug-discovery scientists, synthetic biologists, plant biologists, ecologists, neuroscientists, aging researchers, graduate students, and research program leaders who need to evaluate AI systems against experimental evidence.

The One-Page View

The capability frontier is real and uneven. Structure prediction and selected design tasks are demonstrated. Single-cell representation works; single-cell perturbation prediction is contested. Genome foundation models are early. Systems biology and virtual-organism claims require stronger perturbation evidence. Plant, ecological, neuroscience, and aging AI are now large enough to deserve explicit homes. AI for therapeutics is delivering candidates into trials but not yet approvals. Autonomous laboratories work for bounded tasks. Information hazards require institutional discipline. Reproducibility requires infrastructure investment.

The value in life sciences AI is not the model alone. It is the validation plan, the wet-lab follow-up, the calibration discipline, and the willingness to treat every model output as a hypothesis with explicit uncertainty. Programs that internalise this read the literature differently and produce more reliable science.