Benchmarks for Bio AI

Published

October 7, 2026

Benchmarks are social infrastructure for scientific claims. A good benchmark narrows the space of plausible claims; it does not settle all uses of a model. CASP, CAMEO, PoseBusters, MoleculeNet, scIB, OpenProblems, Therapeutics Data Commons, and the Arc Institute Virtual Cell Challenge each illustrate a different benchmark role: blinded community assessment, continuous server evaluation, validity-aware metrics, shared splits, multi-task evaluation, and an emerging blinded competition. Dangerous-capability evaluations add a second question: whether a model materially lowers the barrier to misuse. The shared discipline is that a leaderboard is a filter, not a validation plan.

Named suites (CASP, CAMEO, PoseBusters, scIB, OpenProblems, TDC, Virtual Cell Challenge) still answer single-task questions. Saez-Rodriguez, Schäfer, Kalavros, and Stolovitzky argue that biomedical foundation models are harder objects: their parameters are framed as parameterized embodiments of the phenomena that generated the training data, so limitation-testing is epistemological as well as leaderboard-shaped (Saez-Rodriguez et al., 2026). Their Nature Methods Perspective pushes three handbook-ready disciplines: ask whether a foundation model can be refuted, verified, or judged mainly by utility; avoid the self-assessment trap where developers define and grade their own finals; and import CASP/DREAM-style independent community referees with withheld evaluation data, shared frameworks, and metric taxonomies aligned to downstream use, because metric choice can reverse perturbation-model headlines. Perspective and evaluation-governance argument, not a new empirical bake-off that ranks today’s single-cell or genomic foundation models.

Learning Objectives

Use this chapter to:

  • Show how benchmarks convert broad AI-biology claims into measurable tasks that can be compared, criticized, and improved.
  • See why a benchmark is only useful if its split, metric, leakage controls, task definition, and biological endpoint match the claim.

Prerequisites: Evaluation Principles for Life Sciences AI for the credibility hierarchy.

Summary: Show how benchmarks convert broad AI-biology claims into measurable tasks that can be compared, criticized, and improved. Some benchmark cultures are mature, especially structure prediction, while newer areas still need better blind tests, prospective evaluation, and biosecurity evaluation that separates capability, uplift, safeguards, operational consequence, and lifecycle governance.

Key point: A benchmark is only useful if its split, metric, leakage controls, task definition, and biological endpoint match the claim. Open question: whether benchmark results predict prospective experimental value rather than leaderboard rank.

Bottom line: Benchmarks connect methods across the handbook by giving molecular, cellular, therapeutic, and automation claims a common evidence discipline.

Field Guide

What is this field trying to solve? Show how benchmarks convert broad AI-biology claims into measurable tasks that can be compared, criticized, and improved.

What is the core idea? A benchmark is only useful if its split, metric, leakage controls, task definition, and biological endpoint match the claim.

What is the current state of the field? Some benchmark cultures are mature, especially structure prediction, while newer areas still need better blind tests and prospective evaluation. Biosecurity evaluation is moving from single-score proxies toward configuration-specific evidence across capability, uplift, safeguards, operational consequence, and lifecycle governance.

What do we know, and what remains open? Known reference points include CASP, CAMEO, PoseBusters, MoleculeNet, Therapeutics Data Commons, scIB, OpenProblems, Virtual Cell Challenge, DOME, WMDP, LAB-Bench, uplift studies, and calibration metrics. What remains open is whether benchmark results predict prospective experimental value rather than leaderboard rank or unsafe capability access.

Why does this matter? Benchmarks connect methods across the handbook by giving molecular, cellular, therapeutic, and automation claims a common evidence discipline.


Introduction

CASP, CAMEO, MoleculeNet, and PoseBusters each illustrate a different benchmark role: blinded community assessment, continuous server evaluation, shared molecular datasets, and physical validity checks (Kryshtafovych et al., 2024; Haas et al., 2018; Wu et al., 2018; Buttenschoen et al., 2024). The single-cell side adds scIB and OpenProblems; the therapeutic-discovery side adds the Therapeutics Data Commons; the cell-perturbation side adds a Virtual Cell Challenge commentary and public results for virtual-cell claims (Roohani et al., 2025; Arc Institute, 2025). AI-bio safety evaluation adds another class: benchmarks and human-uplift studies that test whether a model changes access to hazardous knowledge or practical biological capability.

Demonstrated capability

Demonstrated capability includes benchmark-driven progress in protein structure prediction and increasingly strict evaluation of molecular docking and generation. CASP documents categories beyond single-chain structure, including complexes, RNA, and ligand binding (Kryshtafovych et al., 2024). CAMEO complements biennial CASP cycles with continuous blind server evaluation against newly released protein structures (Haas et al., 2018). PoseBusters demonstrated that physically invalid poses can pass simpler docking metrics (Buttenschoen et al., 2024). GuacaMol standardised goal-directed and distribution-learning benchmarks for de novo molecular design (Brown et al., 2019). scIB and OpenProblems demonstrated community-driven benchmarking for single-cell AI (Luecken et al., 2022; Luecken et al., 2025). Therapeutics Data Commons demonstrated multi-task therapeutic discovery benchmarking (Huang et al., 2022).

For aging and longevity AI, LongevityBench (with Longevity-LLMs and Longevity Claw) is an open evaluation stack aimed at whether current systems can spearhead aging research, complementing clock-centric resources already listed in the aging chapter (Zhavoronkov, Gladyshev et al., 2026). Benchmark suite for research readiness, not proof of healthspan benefit.

Evidence Anchor What It Supports Practical Constraint
CASP and CAMEO Structure prediction assessment Tasks evolve as methods improve
MoleculeNet and TDC Molecular and therapeutic benchmarks Dataset splits shape conclusions
GuacaMol Generative molecule evaluation Benchmark reward functions can be optimised without improving make-test value
PoseBusters Physical validity in docking evaluation One metric can hide failure
scIB and OpenProblems Single-cell community benchmarks Coverage extends as more tasks land
Virtual Cell Challenge Blinded cell-response benchmark Recent; community discipline is emerging
Science sandboxes (MPRAbox, CodonBox) Closed-loop eval of whether agents infer rules or only raise an oracle score Preprint; damp/dry oracles; not a dangerous-capability eval
WMDP Proxy hazardous-knowledge benchmark for biosecurity, cybersecurity, and chemical security Public proxy; not an end-to-end misuse demonstration
LAB-Bench Practical language-agent tasks for biology research Useful biology capability can overlap with dual-use capability
MetagenomicsBench Deterministically graded agentic metagenomics analysis Preprint; scores apply to model-harness configurations, not models alone
Human bio-uplift studies Human-with-model versus baseline evaluation pattern Results depend on the task, baseline, assistance condition, and endpoint

Rao et al. proposed science sandboxes as a closed loop in which an agent chooses experiments, a sealed oracle returns a report, and a lab notebook is scored for whether the agent inferred the rules or only raised a metric (Rao et al., 2026, preprint). In MPRAbox, Claude Opus 4.7 matched or beat 14 human MPRA-library strategies against a Malinois damp oracle; in dry oracles and CodonBox, the same class of agent could max a fitness number without recovering the hidden rule. That is a preprint measurement of scientific reasoning on damp and dry oracles, not a wet-lab discovery result, and a high score is still not proof that the agent understood the system.

MetagenomicsBench applies the same discipline to microbiome analysis. It supplies 100 evaluations built from published metagenomic datasets spanning community structure, host-microbiome associations, microbial function, longitudinal dynamics, and microbiome interventions, each with a deterministic grader that checks whether an agent recovered the original analytical or biological result. Across 10,200 trajectories from 34 model-harness configurations, the strongest configuration passed 60.0%, and greater cost, token usage, and tool use did not consistently track accuracy (Yang et al., 2026, preprint). Trajectory review found agents producing technically plausible analyses while erring on problem interpretation, statistical reasoning, and biological interpretation. Preprint; results are reported per model and harness, so one configuration’s pass rate is not a model ranking.

Biomedical imaging VLMs need the same leakage discipline. MMBU (D’Cunha et al., arXiv 4 June 2026, preprint) spans 35 submodalities from 410 curated datasets and finds that medical fine-tuning wins on PathVQA / VQA-RAD / SLAKE do not reliably survive a broader perception suite: closed-to-open gaps are large, object detection stays near or below random, and adapted vs base models tie on most head-to-heads (arXiv:2606.06696v1). Treat as a biomedical imaging benchmark culture signal, not as a wet-lab or clinical endpoint.

For modality-specific claims that MMBU stress-tests across pathology and microscopy stacks, see Histopathology AI and Microscopy and Cryo-EM AI.

Safety and dangerous-capability evaluation

Beneficial-capability benchmarks ask whether a model helps with a scientific task. Dangerous-capability evaluations ask a different question: whether a model, model scaffold, or tool connection materially lowers the barrier to harmful biological use. The Weapons of Mass Destruction Proxy benchmark is a public proxy for hazardous knowledge in biosecurity, cybersecurity, and chemical security, filtered to avoid directly releasing sensitive operational content (Li et al., 2024, preprint). LAB-Bench measures practical biology-research capabilities such as literature reasoning, protocol planning, database navigation, and sequence manipulation; those are useful research skills, but they become safety-relevant when they overlap with cloning, protocol execution, or agentic lab workflows (Laurent et al., 2024, preprint).

Biosecurity evaluation requires five distinct analytical layers:

  1. Capability measurement: What can the model or agent accomplish under specified elicitation, tool, and access conditions?
  2. Human or agent uplift: How much does system access change performance relative to a matched baseline?
  3. Safeguard effectiveness: Do policy controls restrict hazardous assistance while preserving legitimate scientific use, including under declared adversarial budgets?
  4. Operational validation and risk translation: Does assistance change a consequential real-world bottleneck, rather than only a proxy score?
  5. Governance assurance and lifecycle monitoring: Are controls independently reviewed, monitored after deployment, remediated, and reassessed as systems change?

Threat-model specification and statistical validity cut across all five layers. A result in one layer does not establish the others. A high LAB-Bench score is not an uplift estimate. A refusal rate is not evidence of adversarial robustness. A jailbreak is not proof of operational consequence.

Uplift studies test the marginal effect of model access by comparing performance with and without the model. RAND’s 2024 red-team study found no statistically significant difference in the viability of biological attack plans for the systems and conditions tested (Mouton et al., 2024). OpenAI’s early-warning study found small, non-statistically-significant increases in accuracy and completeness on written biological-threat information-access tasks. It did not test physical construction and presented the result as a starting point rather than a final risk estimate (OpenAI, 2024). A 2026 preprint found substantial novice uplift across bounded digital biology tasks, but did not evaluate physical-laboratory execution or general life-science research productivity (Zhang et al., 2026, preprint). A separate preregistered randomized trial using mid-2025 models found no significant difference in full-workflow completion for novice participants, while secondary analyses remained compatible with a possible modest benefit (Hong et al., 2026, preprint). These studies are time-bounded and do not establish a current capability ceiling.

Current developer assessments address different layers and should not be collapsed into uplift claims. OpenAI reports that it precautionarily treats all three GPT-5.6 models as High biological and chemical capability, while Anthropic reports treating Claude Opus 5 as CB-1 but not CB-2 and applying ASL-3 protections (OpenAI, 2026, system card; Anthropic, 2026, system card). OpenAI describes its evaluations as lower bounds and notes that further wet-lab validation may change its conclusion; Anthropic’s Opus 5 assessment relied on automated chemical and biological evaluations without new expert red-teaming or uplift trials. These are developer-reported capability and safeguard classifications, not independent demonstrations of misuse uplift or operational consequence.

Add Opus 5.5 beside Mythos 5.1: same CB-1 / not CB-2 treatment with expanded bio safeguards, plus the card’s life-sciences capability block (protein design / de novo binders, BioMysteryBench, Protocols) as the current Opus-class developer self-report (Anthropic, 2026, Opus 5.5 system card). Developer-reported RSP/FCF classification and capability scores, not an independent uplift trial.

NIST AI 800-2 is an Initial Public Draft on the validity, transparency, and reproducibility of automated benchmark evaluations. NIST AI 800-3 provides statistical methods for distinguishing performance on a fixed benchmark from performance generalized to a broader task population. Neither is a biosecurity operational standard, but both support explicit estimands, assumptions, uncertainty, and limits on generalization.

For life-sciences teams, the practical rule is direct: do not treat scientific benchmark performance as a release-safety decision. A model can perform well on CASP, TDC, or LAB-Bench and still require separate evaluation for misuse uplift, sensitive protocol assistance, model-weight security, safeguard effectiveness, tool-mediated lab access, and post-deployment monitoring. Safe-proxy TEVV belongs in the benchmark toolkit because it tests capability boundaries without publishing operational recipes for misuse; the information-hazards chapter explains that disclosure layer.

Theoretical capability

Theoretical capability includes prospective discovery benchmarks where models choose experiments and are judged by cost-adjusted learning. This is the right direction for many life sciences tasks, but it is more expensive than static benchmark release. Cost-aware benchmarks that rank methods by expected discovery yield per dollar are emerging in materials and chemistry; broader adoption in biology requires institutional and economic alignment.

Theoretical capability also includes durable bio-uplift benchmarks that stay diagnostic as models become better at tool use, long-horizon planning, and protocol execution. Static multiple-choice tests saturate or leak. Human-uplift studies are slower and require careful safety review. Agentic evaluations need sandboxed tools, safe biological proxies, and preregistered red lines for what cannot be disclosed publicly.

Beyond current capability

Beyond current capabilities includes a universal biological benchmark that ranks all models. Biological tasks differ too much in ground truth, cost, and acceptable error. A benchmark that bridges protein structure, cellular perturbation, drug discovery, and clinical translation in one ranking is incompatible with the heterogeneity of biological evaluation.

Evidence that would change the assessment

Benchmarks become more promising when results predict prospective experimental value, not only leaderboard rank. Stronger benchmark evidence would connect blinded results, biology-aware splits, failure categories, and cost-adjusted learning to the decision the model will change.

Implications for research and program decisions

  • Use benchmarks to reject claims, not only to support them.
  • Prefer splits that reflect intended use.
  • Report failure categories beside average metrics.
  • Hold back prospective tests when the field is likely to overfit public leaderboards.
  • Separate beneficial-capability benchmarks from dangerous-capability evaluations before releasing weights, agents, protocols, or connected lab tools.
  • Use safe-proxy TEVV for dual-use design claims instead of evaluating operationally sensitive sequences or procedures directly.
  • Cite the specific benchmark (CASP15, MoleculeNet, scIB, TDC) with version when version matters.
  • Cross-validate vendor claims against the relevant community benchmark before integrating a tool.