20 Years of Covering The Cancer Genome Atlas
The Enduring Legacy of The Cancer Genome Atlas
By Charlene Lancaster, PhD
Launched in 2006, The Cancer Genome Atlas (TCGA) set out to improve our understanding of cancer by molecularly characterizing numerous human tumors. Teams from 20 institutes across the United States and Canada drove this effort, examining more than 20,000 primary tumor and matched normal samples from 33 cancer types, including lung adenocarcinoma, breast ductal carcinoma, acute myeloid leukemia, and glioblastoma multiforme. This project helped scientists clarify mechanisms that drive cancer development, classify tumors into distinct molecular subtypes, and identify new biomarkers and therapeutic targets. Additionally, the TCGA empowered researchers to perform pan-cancer analyses and uncover shared aberrations across cancer types.
The initiative generated more than 2.5 petabytes of genomic, epigenomic, transcriptomic, and proteomic data, all publicly available to the research community. This open resource continues to support research today, with scientists still using TCGA datasets to explore cancer biology and identify clinically relevant patterns. Since its launch, The Scientist has chronicled TCGA and continues to report on discoveries enabled by this landmark project.
Remembering Daniela Gerhard
Many scientists and administrators played a part in TCGA’s development, including Daniela Gerhard. As a science administrator at the National Cancer Institute for almost two decades, she helped translate TCGA from concept into an operational, multi-institutional program by contributing to early implementation efforts, including sample collection and processing workflows. Beyond TCGA, she was integral in building several other cancer genomics resources. Gerhard passed away unexpectedly in 2021, leaving a legacy closely tied to the foundations of modern cancer genomics research.
How TCGA Transformed Cancer Genomics Research
Anna Barker discussed how The Cancer Genome Atlas project established a new model for studying cancer through shared data and rigorous standards.
Interviewed by Niki Spahich, PhD
Image Caption
Anna Barker, PhD
Chief Strategy Officer
Ellison Medical Institute
Director, Transformative Healthcare Networks
Co-director, Complex Adaptive Systems
Professor, Arizona State University
The Cancer Genome Atlas (TCGA) was conceived at a moment when advances in genome sequencing began to reshape cancer research. Leaders at the National Cancer Institute (NCI) and the National Human Genome Research Institute (NHGRI) developed the project to create a comprehensive, high-quality resource for understanding cancer genomics. By integrating genomic, epigenomic, and clinical data across thousands of samples, TCGA established a blueprint for large-scale, team-based science.
In this interview with The Scientist, Anna Barker, who played a central role in designing and launching the initiative, discussed the successes and challenges associated with TCGA, which led, through a grand collaborative effort, to the characterization of over 20,000 biospecimens spanning 33 tumor types.
How did TCGA come about?
The origins go back to just after the human genome was sequenced. My colleagues and I had undertaken something called the Cancer Genome Project, but it became clear that if we left cancer genome sequencing to individual investigators, it would take many decades to complete. At the same time, technologies were advancing, and there was a growing recognition that biology was becoming increasingly digital.
These things converged: the technology, the experience from the Human Genome Project, and a strong community of scientists capable of working together. That led to what eventually became TCGA.
What challenges did you face when developing TCGA?
There were challenges at every level. We hadn’t done anything like this before. The Human Genome Project had been completed over a long period of time, but this was different.
It was clear that the data coming out of this project would have to be extremely high quality. Up to that point, there wasn’t a high level of rigor in oncology when working with human tissues. So, we organized the pilot project to acquire high-quality tissues with very stringent criteria by standardizing everything—DNA extraction, sample handling, and even how samples were shipped. We centralized a lot of those processes, which had not been done before at that scale.
The other big surprise was the sample quality in existing collections and biobanks. About 30 percent of the samples didn’t meet our standards. That meant that it took years to collect enough usable samples for some cancers.
Also, there was pushback about whether “big science” would take away from individual investigator-driven research. But we showed that there comes a point where big science is necessary to enable small science. TCGA set the stage for that.
The first major TCGA study focused on glioblastoma (GBM). What did that reveal?
Scientifically, we learned a lot. GBM is a deadly cancer, and we were hoping to see what made it so aggressive.1 It has an extraordinary genomic profile because it doesn't have a lot of mutations. But we found that tumors that behaved different clinically also looked different genomically. That allowed us to begin subtyping the disease. We confirmed known pathways and genes, discovered new ones, and showed that large-scale data could replicate and extend findings from smaller studies. It was a very important proof of concept.
What were some key insights from the later pan-cancer studies?
Those studies were about looking across cancers to find commonalities.2 Are there shared pathways or shared genes? They showed the value of having high-quality data. The TCGA dataset had mutations, copy number changes, epigenomics, and clinical data, and one could analyze these in different ways.
Some cancers are driven largely by copy number changes. Others are due to point mutations. So, you start to see common themes and important differences among these cancers. Now we’re in a phase where researchers are integrating TCGA with other datasets, often using AI. That’s where deeper insights are emerging.
What is the legacy of TCGA?
I think it opened up the minds and hearts of scientists to the value of collaboration. It showed that you can do things together that you could never do on your own, and it inspired other science disciplines beyond cancer research because people followed the blueprint of this project. It also set the stage for where we are now in the digital and AI era, showing how to collect data at scale in a way that’s actually useful.
We’re still in the infancy of understanding cancer, but TCGA helped move us forward, turning the big data revolution into something that will ultimately benefit patients. It set the stage for a lot of what we’re doing now, and I’m very proud of this project.
This interview has been condensed and edited for clarity.
References
- McLendon RE. Comprehensive genomic characterization defines human glioblastoma genes and core pathways. Nature. 2008;455(7216):1061-1068.
- Weinstein JN, et al. The Cancer Genome Atlas Pan-Cancer analysis project. Nat Genet. 2013;45(10):1113-1120.
Decoding Cancer’s Regulatory Programs
Andrea Califano explores how gene regulatory networks translate diverse genomic alterations into tumor phenotypes.
Interviewed by Charlene Lancaster, PhD
Image Caption
Andrea Califano, Dr
President, Biohub NY
Professor of Chemical and Systems Biology
Columbia University
Originally trained as a physicist, Andrea Califano brings a different perspective to biology, using quantitative and systems-level approaches to study cancer. As president of Biohub NY and a systems biologist at Columbia University, he focuses on gene regulatory networks to uncover the logic underlying tumor phenotypes.
Previously, The Scientist spoke with Califano about his use of The Cancer Genome Atlas (TCGA) and quantitative algorithms to reconstruct gene regulatory networks and discover new therapeutic targets. In a follow-up interview a decade later, he describes how his group continues to leverage TCGA datasets, together with updated analytical methods, to identify conserved master regulators and tumor regulatory states across cancer types.
What motivates you to study gene regulatory networks?
Because of my background in physics, I tend to focus on the math. That led me to realize there is a fundamental paradox. With 2,000 oncogenes, there are 10400 possible mutation combinations that may cause cancer. By comparison, the universe contains only about 1080 atoms. That means there are far more mutational patterns than atoms in the universe. This suggests scientists cannot defeat cancer by therapeutically targeting individual mutations.
Instead, my team decided to take another approach. If tumor cells look and behave the same despite very different genomic alterations, some integration mechanism must exist to translate those mutations into specific cellular phenotypes. We thought the transcriptional layer was precisely where this integration occurs. We began identifying molecules, including transcription factors (TFs) and cofactors, that control these phenotypes. We call them master regulators and believe that they are better drug targets for defeating cancer.
How do you identify master regulators?
We use Algorithm for the Reconstruction of Accurate Cellular Networks (ARACNe) to assemble gene regulatory networks from large gene expression datasets, such as those available through TCGA.1 This method relies on the data processing inequality theorem to distinguish direct from indirect interactions. Like the game of telephone, where a message becomes distorted as it passes from person to person, the theorem states that information in a noisy system can only be lost but never increase. ARACNe applies this theorem by measuring the mutual information between a TF and a potential target gene. If we find an alternative path that carries more information than the direct TF-target connection, the algorithm classifies the interaction as indirect. Cutting these weakest links leaves us only with the bona fide, direct interactions.
After reconstructing the network, we analyze the expression of genes that ARACNe predicts each TF regulates using Master Regulator Inference Algorithm (MARINa) or Virtual Inference of Protein-activity by Enriched Regulon analysis (VIPER).2,3 This allows us to infer the TF’s activity and determine if it is a master regulator.
What did master regulator analysis reveal about the relationship between genomic alterations and tumor phenotypes?
We used ARACNe and VIPER in our Multiomics Master-Regulator Analysis (MOMA) framework to analyze about 10,000 tumor samples across 20 cancer types from TCGA.4 We identified 112 transcriptionally distinct tumor subtypes defined by master regulators. Despite highly diverse mutational landscapes across cancer types, these subtypes reflect combinations of only a limited number of shared master regulator-driven programs. We found that 24 modules of master regulator proteins help control critical tumor hallmarks, such as DNA repair or mitosis. Combinations of these modules create every tumor subtype. Thus, if we had drugs that target these 24 modules, we could potentially target most tumors.
Why is TCGA a critical resource for your work?
TCGA has been instrumental and foundational to systems biology because it provides a collection of tumor samples with well-defined phenotypes profiled across multiple layers, including genomics, transcriptomics, and epigenomics. We could not have built these gene regulatory networks or identified master regulators without TCGA, as we needed enough samples to characterize each phenotype and capture the full repertoire of cancer states. This large-scale multiomic approach is now inspiring similar efforts in other diseases, such as neurodegenerative disorders, diabetes, obesity, and inflammatory diseases. As a result, an entire range of human cellular phenotypes could not have been systematically studied without TCGA’s seminal contribution.
This interview has been condensed and edited for clarity.
References
1. Margolin AA, et al. ARACNE: An algorithm for the reconstruction of gene regulatory networks in a mammalian cellular context. BMC Bioinformatics. 2006;7(1):S7.
2. Lefebvre C, et al. A human B‐cell interactome identifies MYB and FOXM1 as master regulators of proliferation in germinal centers. Mol Syst Biol. 2010;6(1):MSB201031.
3. Alvarez MJ, et al. Functional characterization of somatic mutations in cancer using network-based inference of protein activity. Nat Genet. 2016;48(8):838-847.
4. Paull EO, et al. A modular master regulator landscape controls cancer transcriptional identity. Cell. 2021;184(2):334-351.e20.
Understanding the Evolution of Viral Oncogenesis Research
Scientists use The Cancer Genome Atlas to deepen their understanding of the role of viruses in triggering cancer.
Interviewed by Iris Kulbatski, PhD
Image Caption
Dirk P. Dittmer, PhD
William R. Kenan, Jr. Distinguished Professor of Microbiology and Immunology
Lineberger Comprehensive Cancer Center
University of North Carolina at Chapel Hill
Twenty years ago, the degree to which certain viruses contribute to cancer was unclear, with some studies questioning the estimates on the percentages of cancer-associated viruses. Certain viruses, however, have been known to be oncogenic for some time. One of these is the Epstein-Barr virus (EBV), which can cause various cancers, including blood cancer.
Dirk Dittmer, a tumor virology researcher at the University of North Carolina at Chapel Hill, studies how human viruses cause tumors and how these cancers can be treated. In an interview with The Scientist, Dittmer described his research on the role of EBV in causing blood cancer, the utility of The Cancer Genome Atlas (TCGA) in this work, and the evolution of viral oncogenesis research.
What is the relationship between viruses and cancers?
One modality is that a virus directly causes cancer, such as human papilloma virus (HPV), which leads to cervical cancer, and EBV, which causes Burkitt lymphoma and is also associated with gastric cancers and nasopharyngeal carcinoma. Cancer is multifactorial, but for these cancers the virus is required. Viruses are present in all of us, and they shape cancer evolution. For instance, EBV is in 90 percent of people. Not everyone who has the virus gets cancer, but everyone’s immune system adapts to its presence in a different way. The immune system is modified with different viral infections, and that, to some degree, underlies how the tumor develops.
In 2018, your team published a study about EBV prevalence in TGCA tumor samples. What were you investigating?
We studied cancers not typically associated with EBV using cancer biopsy samples.1 EBV grows only in B cells, so we asked whether it is present in B cells from patients with cancer and whether these cells differ from those in patients without the virus. We showed that the EBV positive samples had greater B cell diversity. What that means is presently unclear. Even though EBV was associated with an increased and altered immune response in patients with cancer, this is not evidence of causality because we don't have a mechanism to explain why some cancers contain EBV positive B cells and others don’t. It may represent a novel biomarker, which is interesting.
How did you use TCGA in this work?
The RNA in the cancer biopsy samples was sequenced and aligned with the human genome to get a tumor signature. The data is an indirect readout of how many B cells are in the tumor because certain molecules are only expressed on these cells. With lung cancer, for example, a sample will have lung epithelial cells and whatever blood cells happen to be in the lung at that point. Data from these are captured in TCGA, and we used that information to understand the number and types of B cells present in different cancers, and whether they are EBV positive. We can then ask whether there is a difference between EBV positive B cells in brain cancer and EBV negative B cells in breast cancer. EBV is not present in breast cancer tumors, but if you sequence the mixture of cells from the biopsy, you will find human viruses infiltrating them. Such information may help stratify immune therapy patients based on the types of B and T cells in the cancer, rather than on cancer type. We can use TCGA and other data to explore this.
How has the understanding of viral oncogenesis evolved?
Over the last decade we appreciated that certain viruses cause cancer, including HPV, EBV, Merkel cell carcinoma virus, Kaposi sarcoma herpes virus, and hepatitis C virus. These are tabulated by the International Agency for Research on Cancer and are taught in textbooks. Viruses also affect cancer evolution and therapy. We wouldn't have known this without sequencing data to find footprints of these viruses because they are rare events. If the viruses are there, then part of their DNA or RNA is present in TCGA data that allows us, for instance, to look at the presence of HPV in head and neck cancer. Head and neck cancer caused by smoking is different than that caused by HPV, and there are now clinical protocols that use this distinction as a biomarker to determine therapy. The Human Virome program that I'm a part of tries to understand the normal viruses found in people and whether they are different in cancers using TCGA data. Very few cancers are causally associated with viruses, but many are shaped by the presence of viruses.
New technologies deliver qualitatively different and better data. We will learn more from single-cell sequencing or spatial sequencing. That's the hope anyway.
This interview has been condensed and edited for clarity.
Reference
- Selitsky SR, et al. Epstein-Barr virus-positive cancers show altered B-cell clonality. mSystems. 2018;3(5):e00081-18.
Travel Through The Cancer Genome Atlas’ Timeline
Scientists first teamed up for The Cancer Genome Atlas (TCGA) project in 2006, seeking to molecularly characterize mechanisms that drive cancer development and improve diagnostics and treatments. Over the years, this initiative has produced a plethora of publicly accessible genomic, epigenomic, transcriptomic, and proteomic insights, which scientists continue to use and build upon today. The Scientist has chronicled TCGA from its inception and continues to report on discoveries enabled by its findings.
Where It Started | TCGA Today |
Anna Barker played a central role in designing and launching TCGA. After the human genome was sequenced, she and her colleagues began working on the Cancer Genome Project, realizing the need for an all-hands-on-deck collaborative approach. As biological research became increasingly digital and interdisciplinary, new technologies converged with a strong community of scientists capable of working together, leading to the project that eventually became TCGA. | Today, the legacy of TCGA continues to be seen in research across different scientific disciplines, emphasizing the value of collaborative science and large-scale data collection and distribution. According to Barker, TCGA created a blueprint for studies beyond cancer research. Additionally, scientists are now integrating TCGA with other datasets, often using AI, leading to deeper insights and turning the big data revolution into something that will ultimately benefit patients. |
Systems biologist Andrea Califano develops quantitative algorithms to make the most of large gene expression datasets, including TCGA. Starting in 2006, his laboratory has produced tools such as Algorithm for the Reconstruction of Accurate Cellular Networks (ARACNe), Master Regulator Inference Algorithm (MARINa), and Virtual Inference of Protein-activity by Enriched Regulon analysis (VIPER), and applied them to TCGA samples to examine gene regulatory networks across cancer types, seeking out master regulators that govern oncogenic processes. | Califano and his research group continue to leverage TCGA datasets, now implementing updated analytical methods to identify conserved master regulators and tumor regulatory states across cancer types. He credits TCGA as instrumental for his systems biology approach to building gene regulatory networks for master regulator identification, supported by TCGA’s collection of tumor samples with well-defined phenotypes across multiomic layers and cancer states. Today, this approach inspires similar regulatory network profiling in other diseases. |
Although it has not always been clear which and how many cancers can be driven by viral infection, studies using TCGA have helped scientists shed light on virus-associated cancers, including human papilloma virus (HPV) and Epstein-Barr virus (EBV). In 2018, tumor virology researcher Dirk Dittmer investigated cancers that were not typically associated with EBV using TCGA data as an indirect readout of how many EBV-positive B cells were in different tumor samples. He and his team demonstrated that EBV-positive samples had greater B cell diversity, potentially representing a novel biomarker. | Over the last decade, scientists have come to appreciate that certain viruses cause cancer or affect cancer evolution and therapy. Dittmer believes this current knowledge would not exist without sequencing information, including that found in TCGA data. Researchers now know head and neck cancer caused by smoking is different than that caused by HPV, enabling cause-specific clinical protocols to determine optimal therapies for different patients. Today, The Human Virome program uses TCGA data to examine common viruses found in people and whether cancer virome profiles could provide new insights in the future. |