Skip to main content

Genetic Code Expansion with Machine Learning Can Bridge the Lab to Market Gap

Non-canonical amino acids can expand the scope of proteins available for therapeutics and machine learning platforms can accelerate their translation.

Written byTrevor Nicks, PhD
| 5 min read
Amino acids in a long chain are shown on a blue background, representing the expanded genetic code.
Register for free to listen to this article
Listen with Speechify
0:00
5:00

In my first semester of graduate school at Tufts University, I sat across from a young professor as he pitched me on joining his lab to work on genetic code expansion (GCE). Even though I'd just finished a degree in biochemistry, I didn't really grasp what he was talking about. I joined a different lab.

Nine years later, all I can think about is GCE. It will be one of the most important technologies of the 21st century, enabling better medicines, industrial proteins, and new applications of biotechnology. Among all of the potential applications of machine learning and artificial intelligence (AI) in the life sciences, it could be the most exciting.

Genetic Code Expansion is Already Reshaping Medicine and Industry

GCE is the process of engineering biological systems to incorporate “non-canonical” amino acids beyond the 20 that nature typically uses to build proteins. The commercial applications are already here. Many of the popular GLP-1 drugs are made using a form of GCE. Antibody drug conjugates, which are enabling a new generation of more targeted cancer therapies, rely on it too. Beyond therapeutics, incorporating non‑canonical amino acids has been shown to significantly enhance enzyme thermal stability across multiple systems, including large increases in half‑life and resistance to aggregation at elevated temperatures, presenting enormous implications for industrial biotechnology.1

Continue reading below...

Like this story? Sign up for FREE Newsletter updates:

Latest science news storiesTopic-tailored resources and eventsCustomized newsletter content
Subscribe

GCE and related technologies for making proteins and peptides with new chemistries are attracting serious investment and scientific attention globally. Companies like Peptidream in Japan have built substantial businesses on GCE and screening platforms for product discovery. Unnatural Products, which uses chemistry to create peptides with non-canonical amino acids, announced a $45 million Series B investment after securing major deals with a handful of large pharmaceutical companies.

It is clear that GCE is moving from academic novelty toward an industrial platform, but the distance between those two points is where the real challenge—and the real opportunity—lives.

Engineering Complexity Makes GCE an AI Problem

Getting GCE from lab discovery to market-ready product is not straightforward. It requires solving some of the most complex multivariate engineering problems in biology, but that complexity is precisely what makes it one of the most exciting applications of machine learning in the 21st century.

Every non-canonical amino acid added to the genetic code requires engineering two bespoke biomolecules: a new tRNA synthetase and a new tRNA for each additional amino acid. These engineered components must function in concert with the rest of the cell's existing protein production machinery, which is a highly interconnected system with little room for error but limitless space for design variation and chemical diversity.

Researchers are building platforms specifically designed to navigate this space. OrthoRep, which is used in yeast-display systems, offers one approach for engineering tRNA synthetases.2 On the cell-free side, a recent article from researchers at the University of Tokyo demonstrated genetic code expansion that incorporates up to 32 distinct amino acids—maintaining all 20 canonical amino acids while adding 12 non-canonicals—using carefully engineered tRNAs and optimized translation conditions.3 Looking further ahead, engineering ribosomes and other elements of the protein production machinery could unlock even more ambitious implementations of GCE.

The In-House Data Flywheel for GCE and Machine Learning

These new platforms, built to screen and engineer GCE machinery, generate enormous amounts of experimental data. And those data are inherently proprietary.

Unlike the sequence data that populate publicly accessible repositories, this data cannot be downloaded from outside sources, and it cannot be replicated without building the same physical experimental infrastructure. It can only be produced through specific laboratory systems that take years and significant resources to develop.

This is what makes these platforms so strategically valuable for machine learning. AI learns from what worked and what didn't. Proteins either function under manufacturing conditions or they don't. Enzymes either maintain stability or degrade. This tight coupling between prediction and physical reality is exactly the environment where machine learning compounds most effectively.

Take, for example, immunology: Antibody drug conjugates, a rapidly growing class of cancer therapies, depend on the precise chemical attachment of a therapeutic payload to an antibody. The current design still relies heavily on trial and error to optimize this attachment. A platform pairing high-throughput GCE and machine learning analysis could screen thousands of non-canonical amino acid variants in parallel, test their performance against real biological targets, and continuously refine its predictions based on results. This could compress years of development into months or weeks while producing conjugates with properties no conventional approach could have designed.

In industrial biotechnology, enzymes engineered for carbon capture face a different but equally demanding problem: They must function reliably under the high temperatures, pressures, and toxic stress of industrial flue-gas conditions, which rapidly degrade conventional proteins. Screening enzyme variants bearing different non-canonical amino acids under those exact conditions can generate proprietary data on which non-canonical chemistries improve stability. That data exists nowhere else because it can only be produced by building the physical system to test it.

As these models improve on in-house datasets, they accelerate the timeline from lab to market by reducing the trial-and-error testing cycles that currently extend development, making it possible to design cheaper, more complex proteins engineered specifically to survive the conditions of real-world industrial deployment.

AlphaFold for Proteins with Noncanonical Amino Acids

But optimizing GCE machinery is not the only significant opportunity. The real prize is building machine learning models that can predict protein structure and function when noncanonical amino acids are incorporated. Doing so requires in-house data that does not yet exist at the necessary scale.

Consider what made AlphaFold possible. Its success depended on two existing databases: the Protein Data Bank, containing the precise three-dimensional structures of nearly 175,000 proteins, and UniProt, a sequence and function database with over 200 million entries. Decades of experimental work built those resources. AlphaFold found patterns in data that already existed.

For proteins incorporating noncanonical amino acids, however, that foundation does not yet exist. There is no equivalent of the Protein Data Bank for the expanded chemical space GCE unlocks. A database of non-canonical amino acids and proteins built for and with GCE, providing data on the effects on folding, stability enhancements, or degradation, and binding affinities that they display, would provide the necessary training data to rapidly evolve GCE’s place in industrial and academic science.

But that database can only be brought into being by organizations producing the data themselves, in-house and at scale, with the physical infrastructure to generate ground truth.

When that foundation is built, the downstream applications follow: hyper-stable industrial enzymes capable of functioning under conditions that destroy conventional proteins, entirely new classes of drugs with mechanisms unavailable to the standard 20 amino acids, and therapeutics requiring far less frequent dosing because their biological stability has been engineered from the ground up.

GCE is the Bridge from Discovery to Market

What I didn't grasp about GCE as a young graduate student is clear to me now: The potential of GCE to enable new biotechnology products—even whole new product categories—is enormous. So is its complexity, and that complexity is what makes it such an interesting use case for machine learning. There is an enormous opportunity to create new systems that accelerate discovery and enable the manufacturing of products with GCE.

The organizations building the experimental platforms to generate that in-house data are doing more than advancing their own pipelines. They are laying the foundation for one of the most interesting applications of machine learning in biotechnology: making GCE the bridge between a scientific possibility and an actual product with the performance patients and other customers need. The data generated along the way is what makes that bridge possible.

  1. Gurevich A, Levy Y. Non-canonical amino acids as a tool for the thermal stabilization of enzymes. Protein Eng Des Sel. 2023;36(1):1-15.
  2. Furuhata Y, et al. Directed evolution of aminoacyl-tRNA synthetases through in vivo hypermutation. Nat Commun. 2025;16(1):4832.
  3. Katoh T, Suga H. A reprogrammed genetic code consisting of 32 distinct amino acids. Nucleic Acids Research. 2026;54(4):gkag140.
Add The Scientist as a preferred source on Google

Add The Scientist as a preferred Google source to see more of our trusted coverage.

Meet the Author

  • Trevor Nicks smiles in a headshot image.

    Trevor Nicks, PhD is Founder & CEO of Caravel Bio. Trevor completed his PhD in Nikhil Nair’s lab at Tufts University in the Department of Chemical & Biological Engineering. His thesis research focused on developing low-cost tools for improving protein stability and activity for use in cell-free processes. This research laid the foundation for Caravel’s platform.

    View Full Profile

Related Topics

You might also be interested in...
Loading Next Article...
You might also be interested in...
Loading Next Article...
The Scientist Digest cover September 2026
September 2026

Multiplex Microscopy Becomes Easier with Encoded Antibodies

A new system that enables researchers to uniquely tag monoclonal antibodies for use in microscopy could help simplify complex imaging studies.

View this Issue
Rethinking ALS Biomarkers: From Discovery to Clinical Impact

Rethinking ALS Biomarkers: From Discovery to Clinical Impact

Alamar Biosciences logo
Best Practices for qPCR Assay Design and Optimization

Best Practices for qPCR Assay Design and Optimization

Bio-Rad
Beyond the Basics: Strategies for Single-Cell and Spatial Transcriptomics Analysis

Beyond the Basics: Strategies for Single-Cell and Spatial Transcriptomics Analysis

bioxcell
Scientist reviewing cellular and molecular data on a computer in a laboratory.

Building Translation-Ready Biomarkers with Connected Workflows

Danaher Logo

Products

Closeup image of a multi channel pipette dispensing pink liquid into a 96-well plate.

The ASSIST PLUS pipetting robot for affordable workflow automation

Integra Logo
Single cells in suspension

Rapidly isolate primary cells and make uniform single-cell suspensions with Corning® Cell Strainers

Corning logo
Abstract image representing cell membranes linked together.

CellBrite® Steady Membrane Stain: Cell surface staining built for real-time imaging

Biotium
sino biological logo

Monod Bio Licenses AI-designed Protein Technologies to SignalChem Biotech for Custom Discovery Assays