Skip to main content

Rethinking Junk DNA: What Is 98 Percent of the Genome Hiding?

In his book Eclipsed Horizons: Unveiling the Dark Genome, Sudhakaran Prabakaran reports that non-coding DNA produces thousands of peptides linked to health and disease.

Written bySneha Khedkar
| 4 min read
Close-up shot of a magnifying glass examining a molecular DNA strand, signifying poorly-understood parts of the dark genome.
Register for free to listen to this article
Listen with Speechify
0:00
4:00

With completion of the Human Genome Project, researchers uncovered roughly 20,000 genes coding for mRNAs destined to be translated into proteins.1 These account for only about two percent of the entire genome, resulting in the classification of the remaining 98 percent as “junk DNA” that only regulates the protein-coding genes. However, as new data about the genome emerged, some researchers began challenging this convention.

While next-generation sequencing and mass spectrometry approaches helped researchers assign protein-coding genes to certain parts of the genome, “there [were] a lot of other parts of datasets that [we] didn’t know how to make sense of,” said Sudhakaran Prabakaran, a systems and computational biologist at Northeastern University.

Over the past few years, Prabakaran and others applied ribosome profiling, mass spectrometry, and proteogenomics to the genomes of various species and discovered that genes that were conventionally considered non-coding also produce proteins. This is the thesis of Prabakaran’s recent book, Eclipsed Horizons: Unveiling the Dark Genome, wherein he questions the status quo that 98 percent of the genome is junk. He offers evidence indicating that non-coding genomic regions may encode more than 200,000 “dark proteins” associated with health, disease, and biological adaptation.

Tracing Hidden Peptides Across Species

The seeds for this work were sown more than a decade ago when Prabakaran and his colleagues analyzed unannotated transcripts and proteomic data from non-coding RNAs in mouse neurons.2 They found novel translation products that were regulated at the same rate as proteins involved in neuronal activity processes.

A photograph of Sudhakaran Prabakaran wearing a white shirt, grey blazer, and grey cap against a light grey background.

Sudhakaran Prabakaran is a systems and computational biologist at Northeastern University.

Sudhakaran Prabakaran

Over the next few years, Prabakaran and others in the field observed novel open reading frames (nORFs)—genomic loci that produce previously unidentified transcripts and protein products—in a number of species. In his book, he explains how these offer insights into the “dark proteome” encoded by non-coding genes.

Continue reading below...

Like this story? Sign up for FREE Genetics updates:

Latest science news storiesTopic-tailored resources and eventsCustomized newsletter content
Subscribe

For instance, while cichlid fish species share nearly identical protein-coding genes, they display a diverse array of colors and patterns. To understand why, Prahakaran and his colleagues investigated the transcriptomic and proteomic signatures from two cichlid fish species.3

“We showed…[that] the proteins that are being made in the non-coding regions are different for different fishes,” explained Prabakaran. Moreover, evolutionary analyses revealed that the timescale over which the two fish species and the nORFs diverged were similar, indicating that the nORFs could potentially play a role in speciation of cichlid fishes.

Beyond animal models, Prabakaran and his colleagues examined thousands of people’s genomes. Scanning through these, the researchers observed non-protein-coding regions that produced peptides associated with schizophrenia and bipolar disorder.4

Dark Proteins as Building Blocks of Adaptation

By using similar approaches, Prabakaran and others discovered thousands of novel peptides encoded by regions of the human genome previously thought to be non-coding. Researchers recently classified these protein-like molecules as “peptideins.”5

“We [have] curated and cataloged around 250,000 additional sets of proteins. And then we did a lot of analysis on their evolution and their function, dysfunction, structure in different diseases,” explained Prabakaran. According to him, these lines of inquiry revealed that the dark proteome consists of “flexible proteins,” meaning that they can help an organism adapt to changing environments unlike the proteins that arise from typical protein-coding genes.

“These so-called dark proteins coming out of the rest of the 98 percent of the genome are like Lego blocks,” said Prabakaran. New, challenging environments result in dismantling the existing Legos and recreating new structures that help the organism evolve and survive, he explained. “[We show] that new proteins are constantly emerging, and that helps us in adapting.”

Studying the Dark Proteome Could Help Unlock Therapeutic Targets

Prabakaran emphasized that dismissing a majority of the genome as junk DNA was not due to ignorance; it was the logical conclusion one could draw based on what the data indicated by the methods available at the time.

He compared the situation to understanding the universe based on telescope images. The Hubble telescope gave scientists one view. “But now we have the James Webb telescope [which offers a] higher resolution,” said Prabakaran. “We are seeing more galaxies. We are seeing more stars. Now our model of universe is slightly, or maybe completely, different.”

Given the growing evidence that non-coding DNA can lead to proteins, what could be the implications of missing out on the dark proteome? “That's [a] trillion-dollar question,” he said, adding that the current drug discovery efforts revolve around a fraction of the known 20,000 protein-coding genes.

As awareness of the dark proteome increases, Prabakaran hopes that researchers will also mine for novel disease-associated proteins in junk DNA. Wanting to explore a new area of therapeutic discovery and identify drug targets within the dark proteome, Prabhakaran co-founded a company called NonExomics, Inc.

Going forward, he hopes that his book increases awareness about this untouched reserve of proteins for biomedical advances. “In order for us to cure diseases…we have to understand the biology first, and for that we have a better tool now, and that's what the whole concept [of the field] is,” said Prabakaran. “[We want to] try to look holistically for what can be changed to cure a disease.”

Add The Scientist as a preferred source on Google

Add The Scientist as a preferred Google source to see more of our trusted coverage.

Meet the Author

  • Sneha Khedkar

    Sneha Khedkar is an Assistant Editor at The Scientist. She has a Master’s degree in biochemistry, after which she studied the molecular mechanisms of skin stem cell migration during wound healing as a research fellow at the Institute for Stem Cell Science and Regenerative Medicine in Bangalore, India. She has previously written for Scientific American, New Scientist, and Knowable Magazine, among others.

    View Full Profile

Related Topics

You might also be interested in...
Loading Next Article...
You might also be interested in...
Loading Next Article...
The Scientist Digest cover September 2026
September 2026

Multiplex Microscopy Becomes Easier with Encoded Antibodies

A new system that enables researchers to uniquely tag monoclonal antibodies for use in microscopy could help simplify complex imaging studies.

View this Issue
Essential Genes Are Dominantly Activated by Single Transcription Factors

Essential Genes Are Dominantly Activated by Single Transcription Factors

EpiCypher Logo
Rethinking ALS Biomarkers: From Discovery to Clinical Impact

Rethinking ALS Biomarkers: From Discovery to Clinical Impact

Alamar Biosciences logo
Engineering CAR-Neutrophils In Vivo to Target Glioblastoma

Engineering CAR-Neutrophils In Vivo to Target Glioblastoma

Miltenyi
Best Practices for qPCR Assay Design and Optimization

Best Practices for qPCR Assay Design and Optimization

Bio-Rad

Products

Closeup image of a multi channel pipette dispensing pink liquid into a 96-well plate.

The ASSIST PLUS pipetting robot for affordable workflow automation

Integra Logo
Single cells in suspension

Rapidly isolate primary cells and make uniform single-cell suspensions with Corning® Cell Strainers

Corning logo
Abstract image representing cell membranes linked together.

CellBrite® Steady Membrane Stain: Cell surface staining built for real-time imaging

Biotium
sino biological logo

Monod Bio Licenses AI-designed Protein Technologies to SignalChem Biotech for Custom Discovery Assays