How to Read This Handbook
This handbook is organised by research object and workflow. Five parts move from foundations through molecular, cellular, organismal, and therapeutic research. Part VI covers research systems and automation. Part VII covers evaluation, practice, and governance, followed by a closing chapter on emerging frontiers. Each major chapter uses the same field-reference and evidence-utility questions so a reader can locate the evidence level before acting on it. The handbook is a reference, not a tutorial, and does not require cover-to-cover reading.
The shortest paths to handbook value, by topic:
- For model selection and evaluation: Evaluation Principles, then Benchmarks for Bio AI, then the chapter for the specific model class.
- For molecular design work: Protein Structure Prediction, Protein Design and Engineering, Antibody and Biologic Design, with Foundation Models for Biology as background.
- For cells, tissues, and systems biology: Single-Cell Foundation Models, Spatial Omics and Tissue Models, Cell Painting and Image-Based Phenotyping, Perturbation Prediction and Virtual Cells, Systems Biology and Multiscale Modeling.
- For organismal and environmental biology: Neuroscience AI and Brain Foundation Models, Aging and Longevity Biology AI, Plant, Crop, and Agricultural AI, Environmental and Ecological AI, Virtual Organisms and Digital Biology.
- For therapeutic discovery and translation: Target Identification, Small Molecule Generation and ADMET, Chemical Biology and Target Engagement, Cell and Gene Therapy AI, Diagnostics and Biomarker Translation, Clinical Trial AI, Translational Evidence and Failure Modes.
- For automation and autonomous research: Self-Driving Laboratories, Robotic Lab Automation and Cloud Labs, Agentic Science Workflows.
- For hands-on adoption: Toolkit for AI-Augmented Bio Research, then Evaluation Principles and Benchmarks for Bio AI.
- For governance, reproducibility, and institutional readiness: Benchmarks for Bio AI, Reproducibility and Open Science, Information Hazards in Capability Research, Workforce, Compute, and Institutional Readiness.
- For strategic planning: Emerging Frontiers in AI for the Life Sciences, then Workforce, Compute, and Institutional Readiness.
Start With the Decision
A role-specific reading path is useful, but the fastest route often begins with the decision at hand. Start with the decision, then work backward to the biological object, evidence design, and tool. The following map keeps a model name from becoming the organizing principle.
| Decision | Start here | Evidence gate before acting |
|---|---|---|
| Whether a model result deserves laboratory time | Evaluation Principles, then the relevant model-class chapter | Biology-aware split, relevant comparator, uncertainty, and a prespecified validation experiment |
| Whether to adopt a research tool | Toolkit for AI-Augmented Bio Research | Version, license, data boundary, reproducibility record, measured workload, and exit path |
| Whether a generated molecule, protein, or sequence is credible | The corresponding molecular-design chapter, then Benchmarks for Bio AI | Modality-specific validity checks, full candidate denominator, and experimental confirmation |
| Whether a discovery claim is translationally meaningful | Translational Evidence and Failure Modes | Endpoint relevance, external or prospective evidence, comparator, and decision consequence |
| Whether an autonomous workflow is ready for routine use | Agentic Science Workflows and Self-Driving Laboratories | Configured-system evaluation, failure recovery, human oversight, and end-to-end provenance |
| Whether an institutional program is ready to scale | Workforce, Compute, and Institutional Readiness | Named workload, accountable owner, governance, operating cost, and evidence that the workflow improves the existing process |
A benchmark can establish performance on a defined test; it cannot by itself establish scientific utility, operational fit, or translational benefit. Move to the next chapter only after naming which of those claims is actually being evaluated.
Reading Paths by Role
The audience is varied; the question is shared: when does an AI output deserve experimental attention? The shortest path depends on the reader’s role.
Computational biologist
Computational biologists can start with AI for the Life Sciences, continue to Evaluation Principles, and then read the relevant model-class chapters. The critical-evaluation sections show how single-cell baselines, docking validity checks, and biology-aware splits change the weight of headline results.
Biotechnology team lead
Biotechnology team leads can read the Executive Summary, the chapters for model classes on which the program depends, and Workforce, Compute, and Institutional Readiness. The AlphaFold 3 access history and the separately developed Boltz and Chai releases provide a case study in licence, deployment, and reproducibility tradeoffs.
Drug discovery scientist
Drug-discovery scientists can read Target Identification, Small Molecule Generation and ADMET, Clinical Trial AI, and Translational Evidence and Failure Modes in sequence. Rentosertib provides a peer-reviewed case spanning discovery and early clinical development, while its Phase 2a primary endpoint and exploratory outcomes define the claim boundary (Ren et al., 2025; Xu et al., 2025).
Physician-scientist
Physician-scientists can use Variant Effect Prediction for interpretation-framework integration, Clinical Trial AI for regulatory framing, and Translational Evidence and Failure Modes. Protein Structure Prediction and Single-Cell Foundation Models provide the relevant discovery context for drug-design or single-cell work.
Synthetic biologist
Synthetic biologists can read Protein Design and Engineering, Synthetic Biology Design Tools, Self-Driving Laboratories, and Information Hazards in Capability Research. Sequence-generation work should follow applicable institutional review and DNA synthesis order-screening controls.
Plant, ecology, or environmental researcher
Plant, ecology, and environmental researchers can read the relevant organism chapter, then Systems Biology and Multiscale Modeling and Evaluation Principles. The central measurement question is whether a model learns biology or sampling effort, geography, season, or platform artifacts.
Neuroscientist or aging researcher
Neuroscience and aging researchers can read the relevant organism chapter, then Virtual Organisms and Digital Biology and Benchmarks for Bio AI. Neural decoding, brain foundation models, aging clocks, and geroscience models can be useful research instruments without establishing general cognition, healthspan, or organism-level causality.
Graduate student or early-career researcher
Students and early-career researchers can begin with History of AI in the Life Sciences, AI for the Life Sciences, and Evaluation Principles, then use the relevant model-class chapter as a starting reading list.
Research program leader
Research program leaders can read the Executive Summary, AI for the Life Sciences, Workforce, Compute, and Institutional Readiness, and the chapter supporting each material proposal claim. The five-question framework separates demonstrated capability, plausible research, and unsupported aspiration.
The Chapter Reference Standard
Used throughout the handbook. Every major chapter begins by answering the same five reference questions:
- What is this field trying to solve? The biological or research problem that defines the area.
- What is the core idea? The assumptions, constraints, and terms that prevent misreading the field.
- What is the current state of the field? The active methods, model classes, and evidence landscape.
- What do we know, and what remains open? The settled reference points, unresolved questions, and limits of current evidence.
- Why does this matter? The scientific, translational, or program implication.
The Chapter Utility Standard
Used throughout the handbook. Every major chapter answers the same five questions:
- What is demonstrated? Supported by published evidence in peer-reviewed venues, official documentation, or reproducible blind benchmark results. The evidence must be specific to a defined task and dataset.
- What is theoretical? Plausible given current methods but not yet established for routine use. The capability has been shown in selected systems or narrow tasks without proven generalisation.
- What is beyond current capability? Not supported by credible evidence with current systems. The capability is aspirational, has been demonstrated only in toy settings that do not generalise, or requires evidence that has not been produced.
- What would make this more promising? The benchmark, prospective experiment, wet-lab confirmation, clinical endpoint, ecological field test, or independent reproduction that would make the area stronger or reveal its limits.
- What should researchers, biotech teams, funders, and program leaders do with this? The practical consequence for evaluation, adoption, funding, experiment design, or deferral.
If a claim cannot answer these questions, the claim is not specific enough. The point is not to be conservative; it is to be precise. In practice, the evidence standard borrows from blinded community benchmarks such as CASP (Kryshtafovych et al., 2024), biology-specific reporting frameworks such as DOME (Walsh et al., 2021), and validity checks such as PoseBusters (Buttenschoen et al., 2024).
Citation Conventions
Citations are inline hyperlinks. The format depends on source type:
- Peer-reviewed papers:
(Author et al., Year)linking to the DOI. Example:(Jumper et al., 2021)linked to the AlphaFold 2 Nature paper. - Preprints:
(Author et al., Year, preprint)linking to the bioRxiv, medRxiv, or arXiv version. Example:(Zambaldi et al., 2024, preprint)linked to AlphaProteo on arXiv. - Company sources:
(Company, Month Year)linking to the official source. Example:(OpenAI, June 2024)linked to the Color Health partnership announcement. - Government and program pages:
(Org, program page)or(Org, Year)linking to the official page. Example:(ARPA-H IGoR, 2026).
Journal citations are checked as connected author, venue, year, and DOI records. Official and organizational claims are checked against primary pages. A resolving record does not establish that the source supports the sentence; claim-to-source alignment remains part of editorial review.
How to Triage a Capability Claim
The discipline for any new claim (paper, preprint, vendor pitch, conference talk):
- Name the biological object. Sequence, structure, ligand, cell state, tissue, experiment, or clinical endpoint.
- Name the task. Prediction, generation, ranking, classification, or planning.
- Identify the evidence. Peer-reviewed benchmark? Prospective experiment? Self-reported? Independent reproduction?
- Place the claim in a utility question. What is demonstrated, what is theoretical, or what is beyond current capability?
- Name the falsifying experiment. Specify the measurement that would change the assessment.
A claim that survives these five steps is worth evaluating. A claim that fails any of them is not yet ready to inform a program decision.
Cross-Handbook Navigation
This handbook sits in a series. For the clinical, public health, and biosecurity layers of AI in biology, see the Companion Handbooks section on the welcome page. The Life Sciences AI Handbook focuses on the discovery layer; the others focus on their respective downstream layers. Each is self-contained.
A Note on Reading Strategy
The handbook is a reference, not a linear textbook. Most readers should start with the Executive Summary, use the role-specific paths above, then return to individual chapters when a model class, evidence standard, or institutional decision needs review. The Quick Reference appendix consolidates the chapter summaries for scanning before a deeper read.