Advances in AI and long-read sequencing are reshaping how researchers study biology at scale. Yet today’s genomic datasets remain limited in both diversity and completeness, constraining the ability of AI models to generalize across species and uncover new therapeutic opportunities. The Trillion Gene Atlas project, which will collect genomic data from over 100 million previously unstudied species across thousands of global sites, aims to overcome these barriers by generating unprecedented volumes of high-quality genomic data.
In this Innovation Spotlight, Glen Gowers of Basecamp Research and Christian Henry of PacBio discuss the scientific vision behind the Trillion Gene Atlas, the role of AI in therapeutic design, and how large-scale genomic sequencing could transform drug discovery and molecular biology.
What was the inspiration for the Trillion Gene Atlas project, and what do you hope to achieve?

Glen Gowers, PhD
Co-founder and CEO
Basecamp Research
Gowers: We released our EDEN foundation model in January 2026. It was trained on 10 trillion nucleotides of totally new evolutionary data, and since this represented about an order of magnitude more data than is publicly available, we were able to establish new scaling laws that show how generalized AI model “intelligence” scales with dataset size. Aside from the technical detail, this gives us a clear blueprint for improving future AI models by orders of magnitude: The answer is data scaling, and this is exactly what we designed the Trillion Gene Atlas to do. At a trillion genes, it represents about three orders of magnitude or 1,000 times more training data than is available in public data and so will push AI models into a new territory of performance.1
Henry: With the Trillion Gene Atlas, we hope to prove that long-read sequencing is no longer confined to specialist use cases. The technology can now scale to deliver complete, information-rich genomes that capture the full complexity of biology at industrial levels. The project’s results should demonstrate that long-read accuracy is non-negotiable for building reliable AI models and improving how we understand disease and design new therapeutics.
Why is studying new species so promising for drug discovery?
Gowers: We’re in a new era of machine learning whereby foundation models are exceptionally good at specific tasks, without the cost of losing generalizability. Key to this success has been first teaching these foundation models very broad spectrum rules of language design. Just as this was done using the internet’s text data for models like ChatGPT, in biology we must first teach a foundation model the language of DNA, the most fundamental language of life, across an enormously diverse range of species. Only then can we point these foundation models to specific tasks, clinical or otherwise. In this way, we will be able to benefit from both the depth and breadth that foundation models have afforded other fields. To date, the breadth component of training data has been severely lacking, with almost 70 percent of all public data coming from just five species. Unless we can get through this systemic data bias, foundation models will suffer. The Trillion Gene Atlas is addressing this bottleneck head on.
What are the key gaps in current genomic datasets, and how do they constrain today’s AI models?
Gowers: Today’s genomic datasets are heavily skewed toward a tiny fraction of life on Earth. Imagine if ChatGPT were trained on just five books. As a result, most AI models today don’t really understand biology yet. They’re good at spotting patterns they’ve already seen, but they struggle the moment you move into something new, such as a novel protein, a rare organism, or a different environment. The data is the bottleneck for the models. If we want AI that can generalize across biology, we need to train it on the full diversity of life, not just the narrow slice we’ve explored so far.

Christian Henry
President and CEO
PacBio
Henry: Today’s genomic datasets are limited not just by missing millions of species, but by the quality of the data scientists currently have. Many species’ genomes remain incomplete, with complex regions, tandem repeats, and highly homologous genes missed or mischaracterized by traditional sequencing approaches. This means AI models are often trained on incomplete or distorted representations of biology. PacBio’s role in this project is to deliver the accuracy and completeness needed to build reliable and trustworthy AI models.
What advantages do high-accuracy long reads offer when working with complex environmental samples compared to traditional sequencing approaches?
Henry: Traditional sequencing approaches fragment DNA into short reads. This makes it difficult to reconstruct genomes and resolve challenging regions, especially in samples where many organisms are closely related, such as soil microbes. In contrast, PacBio HiFi Sequencing long-reads preserve long stretches of native DNA, retaining genomic context and enabling accurate genome assembly, improved detection of structural variation, and clear distinction between highly similar sequences, down to the subspecies or strain level. In an AI project, that level of resolution is critical to ensure the data used to train models reflects the true complexity of biology rather than a fragmented approximation.
What were the major technological breakthroughs that made it feasible to sequence trillions of genes?
Henry: PacBio has made significant progress in increasing throughput and reducing cost per genome, moving long-reads from a specialist tool to operating at industrial scale. Improvements across the broader workflow, from sample processing to computing, have also removed analysis bottlenecks. These advances mean it’s now feasible to generate the volume and quality of data needed to power initiatives like the Trillion Gene Atlas.
What is AI’s role in this project? How can you design therapeutics directly from a disease prompt?
Gowers: AI is central. The technology both depends on vast, diverse datasets and is the only way to make sense of them at scale. By training foundation models on genomic data, we can learn the underlying rules of biology rather than simply cataloguing it. Designing therapeutics from a disease prompt means translating a biological problem, such as a disrupted pathway, into a functional solution. Instead of screening existing compounds, the model can generate novel proteins, peptides, or genetic designs informed by evolution. With partners like NVIDIA and Anthropic, we can scale the computing and reasoning needed to make this practical.
How broad is the potential impact of this dataset, and where do you expect to see the earliest breakthroughs?
Gowers: There are tens of thousands of genetic diseases today where we know exactly what’s broken, we know the gene, but we still don’t have a way of programming a medicine to fix it. By training on much more of the diversity of life, we can move from just predicting what biology does to actually designing new biological solutions. The first place this shows up will be in drug development, but the same approach applies much more broadly, from industrial biotech to environmental systems.
Henry: We expect the earliest breakthroughs to be in areas where challenging biology has held progress back, such as rare disease and oncology. Previous long-read-based studies, such as those by our HiFi Solves Consortium, have already shown that resolving complex variation can unlock new disease insights. At greater scale, those gains could translate into more confident target discovery.
How do you see the kind of interdisciplinary collaboration involved in the Trillion Gene Atlas reshaping the future of molecular biology and drug discovery over the next decade?
Henry: The project reflects a growing convergence between high-quality genomic data and AI. As models become more central to drug discovery, the need for information-rich sequencing is critical. We’re proud to be part of some of the first initiatives helping to bring together genomics leaders, AI innovators, and pharma. Advances in sequencing and computation will work hand in hand to accelerate drug discovery and deepen our understanding of disease.
Gowers: A big shift happening in biology is from describing what exists and tweaking around the edges to actually predicting and designing new systems with intention and high predictability. At Basecamp, that comes down to pushing biological AI models to the absolute forefront of the AI field, bringing together the latest research and tools to bear on this problem. Partnerships are key to this, the opportunity here is greater than any one company, and finding the right models to work together effectively is critical to undergo the fastest possible rate of progress. That’s what starts to turn biology into something we can reason about and engineer, and ultimately redesign how we approach drug development.
- Munsamy G, et al. Designing AI-programmable therapeutics with the EDEN family of foundation models. bioRxiv. 2026. Accessed May 22, 2026.


















