Billions of dollars are spent annually to apply artificial intelligence (AI) to life sciences and drug discovery.1 However, we have yet to see a flood of new drugs and therapies on the market that can be attributed to AI, despite its tremendous potential to accelerate discovery. The root cause of this productivity gap comes down to data.
First, the models with the greatest potential to transform research, such as virtual cell models, are trained on biological data. Compared with the massive datasets used to train large generalist language models like ChatGPT and Claude, the data available for training AI models in biology is minuscule.
Further, we, as researchers, don’t trust the data in our databases and published literature. Nearly three in four biomedical researchers believe our field is gripped by a reproducibility crisis.2 The National Institutes of Health (NIH) has taken notice, launching an initiative to elevate replication and reproducibility as foundational to "gold standard science."
How can AI solve drug discovery when the experts don’t trust the training data?
The solution isn’t to slow down model development and pull focus back to the wet lab. Instead, we have an opportunity to leverage AI to plan and generate high-quality datasets that are useful to models and researchers. The rise of AI is the single biggest opportunity we have to fix biology's reproducibility problem by improving our data practices to make better models possible.
The Protein Data Bank (PDB) is the Exception That Proves the Rule
Protein structure prediction models, such as AlphaFold, have succeeded despite this data problem because of the types of data they rely on.3 The PDB is not only highly organized, but the data it contains is highly reliable.
Most other types of training data in biology, particularly -omics data, are difficult to replicate without access to identical physical starting materials, equipment, and methods. Even if these issues were resolved through adherence to new standards, it could take decades at current rates to populate new databases with datasets approaching the quality of the PDB.

AI models are only as reliable as the biological data they are trained on—making authenticated, standardized datasets the foundation of AI-driven discovery.
ATCC
Gold-standard science requires authenticated biological materials. AI models benefit from data aggregation, meaning that many datasets on the same biological materials should be combined for model training.
But cell lines, microbial strains, and organoids that researchers study are physical, living things. In practice, biological materials may not be what they seem. Cell line genomes, for example, are often unstable, and passaging often leads to mutations and genomic rearrangements. Researchers attempting to replicate studies using their own copies of the same cell lines may not realize that the digital sequence data can diverge significantly from the actual cell line tested.
The variability of physical biological samples incurs real costs. Consider the case of HeLa cell contamination in commonly used research cell lines such as HEp-2 and INT 407. A 2021 analysis found that these two false cell lines alone appear in nearly 10,000 published articles.4 Assuming an average of five citations per article, more than $4.9 billion may have been spent supporting research based on these two unauthenticated cell lines, with costs potentially reaching $14.8 billion under a more inclusive estimate. The human cost of delayed discoveries and the downstream erosion of scientific trust compounds these figures further.
Metadata amplifies the problem. A 2021 analysis published in Clinical Infectious Diseases found that more than a quarter of foodborne microbiological samples in public sequence databases were missing key metadata attributes.5 Without standardized, complete metadata, researchers cannot reliably compare datasets across studies, assess methods, or verify that two experiments actually tested the same thing. AI models trained on such datasets risk encoding and propagating the very inconsistencies researchers are trying to overcome.
In drug development, this matters enormously. R&D costs now exceed $3.5 billion per novel drug, reflecting a five-decade decline in pharmaceutical R&D efficiency.6 Part of that productivity gap reflects the compounding inefficiency of building on unvalidated biological assumptions.
The Fix Requires Stronger Data Infrastructure
Solving the digital biology challenge demands a stronger infrastructure in which digital data is anchored to authenticated biological materials, supported by standardized metadata, and governed by interoperable frameworks that ensure traceability from insights back to their sources.
For researchers, this means insisting that every dataset they work with can answer three fundamental questions:
- Where did this data come from?
- How was it generated and validated?
- Can it be traced back to a known, authenticated biological source?
These cannot merely be “check-the-box” requirements. They must serve as the conditions under which data can be trusted.
Organizations that develop and maintain authenticated biological reference collections are uniquely positioned to model what this looks like in practice. Biobanks and curated repositories form an interconnected ecosystem critical to establishing practical standards for material and data sharing, making it easier to collaborate across organizations. Researchers should demand the reliability that stems from biobanks, repositories, and service providers improving interoperability.
AI Can Help Us Solve Our Data Problem
AI models and human scientists trade in different currencies. While data generation is at the core of what we do, we package this data into readable stories and figures. Raw data, the software pipelines used to generate it, and the detailed methods needed to reproduce it are often buried in supplemental materials, if included at all.
Large datasets may be uploaded to public repositories such as Genbank, but they often lack critical metadata, are misannotated, and cannot be matched to their original source material. The infrastructure we’ve built to serve information and data to scientists doesn’t provide what AI models actually need: well-curated, labeled data that is consistent across studies.
AI, automation, and creative scientists can accelerate data acquisition and model training if we build with authenticated physical materials. Imagine being able to start work at two different contract research organizations or cloud labs with the same cell line without shipping cells to each provider or downloading a protocol for your favorite microbial strain and having it work on the first try.
If multiple datasets are published on your favorite cell line, you should trust that they all share the same parental genome sequence. For developers of AI models, such as virtual cell models, datasets derived from the same physical materials should be easier to aggregate for model training. Biology is fantastically complex, but cell-to-cell variability is one variable that can be controlled. Proving out new scaling laws for biological datasets could create a virtuous cycle of standards adoption and improved model performance.
The Data-Driven Future
Biological data is hard won; anyone who has picked up a pipette can appreciate the effort that goes into conceiving, conducting, and verifying the results of biological experiments. For much of the 21st century, the cost of DNA sequencing has improved faster than Moore’s law, yet the cost of drug discovery has moved in the opposite direction.7 We may be able to reverse that trend if we can make every dataset count.
Many inflection points in the history of biology have arisen from the influx of new viewpoints that changed the practice and language of the science. The arrival of AI in biology is changing how we talk about and do biology. “Non-computational” biologists are now learning to write code and conduct complex bioinformatic analysis with the assistance of AI. AI specialists are applying what they’ve learned to biology for the first time. This inflection point offers a unique opportunity to elevate our digital and physical material practices to accelerate progress.
Any strides we make now in closing the digital-physical divide for biological data will pay off for decades to come. Ideally, these improvements in data quality will provide more immediate feedback through improved model performance and accelerated learning. The next generation of AI-driven discoveries, from CRISPR applications to virtual cell models to personalized therapies for rare diseases, will require datasets that can be trusted as deeply as the physical materials from which they originate.
- Artificial Intelligence in Drug Discovery Market. MarketsandMarkets Research. 2024.
- Cobey KD, et al. Biomedical researchers’ perspectives on the reproducibility of research. PLOS Biology. 2024;22(11):e3002870.
- Jumper J, et al. Highly accurate protein structure prediction with AlphaFold. Nature. 2021;596(7873):583-589.
- Korch CT, Capes-Davis A. The extensive and expensive impacts of HEp-2 [HeLa], Intestine 407 [HeLa], and other false cell lines in journal publications. SLAS Discovery. 2021;26(10):1268-1279.
- Pettengill JB, et al. Interpretative labor and the bane of non-standardized metadata in public health surveillance and food safety. Clin Infect Dis. 2021;73(8):1537-1539.
- Fernald KDS, et al. The pharmaceutical productivity gap: incremental decline in R&D efficiency despite transient improvements. Drug Discov Today. 2024;29(11):104160.
- Scannell JW, et al. Diagnosing the decline in pharmaceutical R&D efficiency. Nat Rev Drug Discov. 2012;11(3):191-200.


















