Expanding the Genetic Alphabet

All life on Earth relies on a genetic code built from four building blocks. Researchers at the University of California San Diego have now shown that one of the most vital enzymes in biology can read and transcribe an expanded, eight-letter genetic alphabet. This discovery suggests that cells possess the molecular machinery needed to process synthetic genetic information. It marks a significant step toward the goal of creating custom biological systems that produce compounds not found in nature.

The research focused on RNA polymerase. This enzyme reads DNA and creates RNA, which is the initial phase of gene expression. The team used biochemical experiments alongside high-resolution cryo-electron microscopy to view the process. These techniques allowed them to capture structural snapshots at a scale smaller than a single atom. They observed RNA polymerase from E. coli bacteria as it recognized and incorporated two synthetic base pairs. These letters exist outside the standard natural set. The snapshots confirmed that the enzyme processes synthetic DNA letters using the same structural and biochemical signals as natural base pairs. This helps explain how such an expanded alphabet can be accurately transcribed.

Insights from the PNAS Study

In a related effort published in PNAS, the same research team discovered that RNA polymerase also recognizes another pair of synthetic bases. These bases do not rely on hydrogen bonds to remain stable. The ability of the enzyme to work with these non-natural components shows that biological systems are more flexible than scientists previously assumed. The findings were detailed in a study led by Dong Wang, a professor at the UC San Diego Skaggs School of Pharmacy and Pharmaceutical Sciences.

This study, titled 'Structural Basis of Transcription of the Hachimoji Eight-Letter Alphabet by E. coli RNA Polymerase', appeared in Nature Communications on September 2, 2026. The PNAS work, 'Hydrophobic unnatural base pair promotes trigger loop closure and catalysis in cellular RNA polymerase independent of hydrogen bonding', was published on August 12, 2026. Both papers highlight the mechanics behind synthetic transcription. They provide a structural foundation for how natural enzymes can tolerate modifications to the genetic alphabet.

Future Implications for Synthetic Biology

This work has practical value beyond the basic study of biological mechanisms. Previous research has already used expanded genetic alphabets to develop synthetic DNA molecules capable of identifying liver cancer cells. By revealing how RNA polymerase processes these non-natural letters, the team has built a foundation for future technologies. These advancements could lead to novel diagnostics, new classes of therapeutics, and specifically engineered biological organisms.

Scientists now look toward creating biological systems that serve specific industrial or medical purposes. The ability to manipulate the fundamental language of DNA opens doors to producing drugs or materials that are currently impossible to synthesize. While these studies focus on E. coli as a model, the broader implication is that the core enzyme machinery of life is compatible with expanded codes. Future researchers must now determine how these synthetic systems function within more complex, living organisms. This discovery provides the roadmap for such work.