Mapping the Regulatory Language of DNA

The human genome consists of roughly 3 billion DNA letters. This vast sequence contains over 9 billion potential single-letter variations. Testing the biological impact of each change in a laboratory setting remains impossible due to the sheer scale of the task. Google DeepMind’s new AlphaGenome Atlas provides a searchable, comprehensive resource designed to help scientists parse this data. Available starting today, the project cataloghs molecular-effect predictions for billions of genetic changes.

Stowers Institute for Medical Research investigator Julia Zeitlinger, Ph.D., collaborated with Google DeepMind’s science team to develop the tool. Led by Vice President of Science Žiga Avsec, Ph.D., the team mapped the specific DNA patterns that regulate biological processes within cells. These regulatory motifs act as instructions for gene activity, determining when and how genes function. By analyzing these signals, researchers can now begin to decode the regulatory language of the human genome.

Advancing Computational Biology Through Collaboration

Zeitlinger and her lab have long been active in computational biology. In 2019, she collaborated with Avsec to develop BPNet, an artificial intelligence framework used to analyze genome-wide data. More recently, her team introduced PISA, a method for generating high-resolution visualizations of what AI models learn from DNA sequences. The AlphaGenome Atlas integrates this deep biological knowledge with the predictive power of advanced AI models.

Melanie Weilert, a bioinformatics scientist in the Zeitlinger lab and a lead author on the project, played a key role in building the resource. The Atlas addresses one of the fundamental questions in biology: how cells determine which genes to activate and which to repress. Because different cell types use different regulatory languages, creating a universal map of these processes has remained a difficult challenge for decades.

Implications for Disease Research and Discovery

AlphaGenome Atlas provides thousands of molecular-effect predictions for every variant across hundreds of tissues. These predictions form the basis of the new AlphaGenome Variant Impact, or AVI, score. This measure helps researchers rank genetic variants by their potential impact on protein-coding and non-coding regions. It integrates predictions from the Atlas with evolutionary conservation data to pinpoint biological mechanisms.

Research institutions are already using the tool to solve practical problems. Scientists at the Broad Institute used the AVI score to identify a previously overlooked non-coding variant linked to an unsolved rare disease case. Simultaneously, researchers at the University of Exeter examined genomic data from 54,000 UK Biobank participants to find new associations between rare non-coding variants and protein levels. This illustrates how the Atlas can shift research from broad exploration to focused investigation.

While the Atlas is a powerful tool, it does not replace laboratory work. It serves as a guide for researchers to prioritize the most promising variants and biological questions. This approach helps focus finite experimental resources on the most impactful leads. By providing public access through a web browser, the team hopes to lower the barrier for scientists to investigate genomic variation at scale. The full findings and the resource are now available to the scientific community on bioRxV.