Skip to main content

AI for Drug Discovery Needs Trustworthy Biological Data

To unleash the potential of AI in life sciences and drug discovery, biological datasets must be authenticated, validated, and supported by standardized metadata.

Written byPatrick Boyle, PhD
| 5 min read
Human skin fibroblast cells are shown via immunofluorescence, stained cyan, blue, and red.
Register for free to listen to this article
Listen with Speechify
0:00
5:00

Billions of dollars are spent annually to apply artificial intelligence (AI) to life sciences and drug discovery.1 However, we have yet to see a flood of new drugs and therapies on the market that can be attributed to AI, despite its tremendous potential to accelerate discovery. The root cause of this productivity gap comes down to data.

First, the models with the greatest potential to transform research, such as virtual cell models, are trained on biological data. Compared with the massive datasets used to train large generalist language models like ChatGPT and Claude, the data available for training AI models in biology is minuscule.

Further, we, as researchers, don’t trust the data in our databases and published literature. Nearly three in four biomedical researchers believe our field is gripped by a reproducibility crisis.2 The National Institutes of Health (NIH) has taken notice, launching an initiative to elevate replication and reproducibility as foundational to "gold standard science."

How can AI solve drug discovery when the experts don’t trust the training data?

The solution isn’t to slow down model development and pull focus back to the wet lab. Instead, we have an opportunity to leverage AI to plan and generate high-quality datasets that are useful to models and researchers. The rise of AI is the single biggest opportunity we have to fix biology's reproducibility problem by improving our data practices to make better models possible.

Continue reading below...

Like this story? Sign up for FREE Newsletter updates:

Latest science news storiesTopic-tailored resources and eventsCustomized newsletter content
Subscribe

The Protein Data Bank (PDB) is the Exception That Proves the Rule

Protein structure prediction models, such as AlphaFold, have succeeded despite this data problem because of the types of data they rely on.3 The PDB is not only highly organized, but the data it contains is highly reliable.

Most other types of training data in biology, particularly -omics data, are difficult to replicate without access to identical physical starting materials, equipment, and methods. Even if these issues were resolved through adherence to new standards, it could take decades at current rates to populate new databases with datasets approaching the quality of the PDB.

Scientist wearing safety glasses sitting in front of a computer.

AI models are only as reliable as the biological data they are trained on—making authenticated, standardized datasets the foundation of AI-driven discovery.

ATCC

Gold-standard science requires authenticated biological materials. AI models benefit from data aggregation, meaning that many datasets on the same biological materials should be combined for model training.

But cell lines, microbial strains, and organoids that researchers study are physical, living things. In practice, biological materials may not be what they seem. Cell line genomes, for example, are often unstable, and passaging often leads to mutations and genomic rearrangements. Researchers attempting to replicate studies using their own copies of the same cell lines may not realize that the digital sequence data can diverge significantly from the actual cell line tested.

The variability of physical biological samples incurs real costs. Consider the case of HeLa cell contamination in commonly used research cell lines such as HEp-2 and INT 407. A 2021 analysis found that these two false cell lines alone appear in nearly 10,000 published articles.4 Assuming an average of five citations per article, more than $4.9 billion may have been spent supporting research based on these two unauthenticated cell lines, with costs potentially reaching $14.8 billion under a more inclusive estimate. The human cost of delayed discoveries and the downstream erosion of scientific trust compounds these figures further.

Metadata amplifies the problem. A 2021 analysis published in Clinical Infectious Diseases found that more than a quarter of foodborne microbiological samples in public sequence databases were missing key metadata attributes.5 Without standardized, complete metadata, researchers cannot reliably compare datasets across studies, assess methods, or verify that two experiments actually tested the same thing. AI models trained on such datasets risk encoding and propagating the very inconsistencies researchers are trying to overcome.

In drug development, this matters enormously. R&D costs now exceed $3.5 billion per novel drug, reflecting a five-decade decline in pharmaceutical R&D efficiency.6 Part of that productivity gap reflects the compounding inefficiency of building on unvalidated biological assumptions.

The Fix Requires Stronger Data Infrastructure

Solving the digital biology challenge demands a stronger infrastructure in which digital data is anchored to authenticated biological materials, supported by standardized metadata, and governed by interoperable frameworks that ensure traceability from insights back to their sources.

For researchers, this means insisting that every dataset they work with can answer three fundamental questions:

  • Where did this data come from?
  • How was it generated and validated?
  • Can it be traced back to a known, authenticated biological source?

These cannot merely be “check-the-box” requirements. They must serve as the conditions under which data can be trusted.

Organizations that develop and maintain authenticated biological reference collections are uniquely positioned to model what this looks like in practice. Biobanks and curated repositories form an interconnected ecosystem critical to establishing practical standards for material and data sharing, making it easier to collaborate across organizations. Researchers should demand the reliability that stems from biobanks, repositories, and service providers improving interoperability.

AI Can Help Us Solve Our Data Problem

AI models and human scientists trade in different currencies. While data generation is at the core of what we do, we package this data into readable stories and figures. Raw data, the software pipelines used to generate it, and the detailed methods needed to reproduce it are often buried in supplemental materials, if included at all.

Large datasets may be uploaded to public repositories such as Genbank, but they often lack critical metadata, are misannotated, and cannot be matched to their original source material. The infrastructure we’ve built to serve information and data to scientists doesn’t provide what AI models actually need: well-curated, labeled data that is consistent across studies.

AI, automation, and creative scientists can accelerate data acquisition and model training if we build with authenticated physical materials. Imagine being able to start work at two different contract research organizations or cloud labs with the same cell line without shipping cells to each provider or downloading a protocol for your favorite microbial strain and having it work on the first try.

If multiple datasets are published on your favorite cell line, you should trust that they all share the same parental genome sequence. For developers of AI models, such as virtual cell models, datasets derived from the same physical materials should be easier to aggregate for model training. Biology is fantastically complex, but cell-to-cell variability is one variable that can be controlled. Proving out new scaling laws for biological datasets could create a virtuous cycle of standards adoption and improved model performance.

The Data-Driven Future

Biological data is hard won; anyone who has picked up a pipette can appreciate the effort that goes into conceiving, conducting, and verifying the results of biological experiments. For much of the 21st century, the cost of DNA sequencing has improved faster than Moore’s law, yet the cost of drug discovery has moved in the opposite direction.7 We may be able to reverse that trend if we can make every dataset count.

Many inflection points in the history of biology have arisen from the influx of new viewpoints that changed the practice and language of the science. The arrival of AI in biology is changing how we talk about and do biology. “Non-computational” biologists are now learning to write code and conduct complex bioinformatic analysis with the assistance of AI. AI specialists are applying what they’ve learned to biology for the first time. This inflection point offers a unique opportunity to elevate our digital and physical material practices to accelerate progress.

Any strides we make now in closing the digital-physical divide for biological data will pay off for decades to come. Ideally, these improvements in data quality will provide more immediate feedback through improved model performance and accelerated learning. The next generation of AI-driven discoveries, from CRISPR applications to virtual cell models to personalized therapies for rare diseases, will require datasets that can be trusted as deeply as the physical materials from which they originate.

  1. Artificial Intelligence in Drug Discovery Market. MarketsandMarkets Research. 2024.
  2. Cobey KD, et al. Biomedical researchers’ perspectives on the reproducibility of research. PLOS Biology. 2024;22(11):e3002870.
  3. Jumper J, et al. Highly accurate protein structure prediction with AlphaFold. Nature. 2021;596(7873):583-589.
  4. Korch CT, Capes-Davis A. The extensive and expensive impacts of HEp-2 [HeLa], Intestine 407 [HeLa], and other false cell lines in journal publications. SLAS Discovery. 2021;26(10):1268-1279.
  5. Pettengill JB, et al. Interpretative labor and the bane of non-standardized metadata in public health surveillance and food safety. Clin Infect Dis. 2021;73(8):1537-1539.
  6. Fernald KDS, et al. The pharmaceutical productivity gap: incremental decline in R&D efficiency despite transient improvements. Drug Discov Today. 2024;29(11):104160.
  7. Scannell JW, et al. Diagnosing the decline in pharmaceutical R&D efficiency. Nat Rev Drug Discov. 2012;11(3):191-200.
Add The Scientist as a preferred source on Google

Add The Scientist as a preferred Google source to see more of our trusted coverage.

Meet the Author

  • Patrick Boyle wears a dark jacket while standing in front of a brick wall.

    Patrick Boyle is Interim Chief Scientific Officer at ATCC and a Founding Partner at American Wetware, a design studio for biology. A bioengineer with more than 20 years of experience, he applies design, computation, and automation to commercialize biology and advises companies on scientific strategy.

    Patrick was an early employee at Ginkgo Bioworks, where he helped design and scale the company's core platform for engineering organisms—turning foundational research into commercial products across pharmaceuticals, agriculture, and industrial biotechnology. As Ginkgo's first Chief Scientific Officer, he led teams that shipped production biological systems for Fortune 500 partners. 

    He was a fellow of the Johns Hopkins Emerging Leaders in Biosecurity Program and served on the Board on Life Sciences at the National Academies of Sciences, Engineering, and Medicine. He holds a PhD from Harvard Medical School and a Bachelor of Science degree from MIT.

    View Full Profile

Related Topics

You might also be interested in...
Loading Next Article...
You might also be interested in...
Loading Next Article...
The Scientist Digest cover September 2026
September 2026

Multiplex Microscopy Becomes Easier with Encoded Antibodies

A new system that enables researchers to uniquely tag monoclonal antibodies for use in microscopy could help simplify complex imaging studies.

View this Issue
Essential Genes Are Dominantly Activated by Single Transcription Factors

Essential Genes Are Dominantly Activated by Single Transcription Factors

EpiCypher Logo
Rethinking ALS Biomarkers: From Discovery to Clinical Impact

Rethinking ALS Biomarkers: From Discovery to Clinical Impact

Alamar Biosciences logo
Engineering CAR-Neutrophils In Vivo to Target Glioblastoma

Engineering CAR-Neutrophils In Vivo to Target Glioblastoma

Miltenyi
Best Practices for qPCR Assay Design and Optimization

Best Practices for qPCR Assay Design and Optimization

Bio-Rad

Products

Closeup image of a multi channel pipette dispensing pink liquid into a 96-well plate.

The ASSIST PLUS pipetting robot for affordable workflow automation

Integra Logo
Single cells in suspension

Rapidly isolate primary cells and make uniform single-cell suspensions with Corning® Cell Strainers

Corning logo
Abstract image representing cell membranes linked together.

CellBrite® Steady Membrane Stain: Cell surface staining built for real-time imaging

Biotium
sino biological logo

Monod Bio Licenses AI-designed Protein Technologies to SignalChem Biotech for Custom Discovery Assays