AI Decodes DNA ‘Initiator’ Sequences in Human Genes

August 26, 2026 — A research team led by Professor James T. Kadonaga at the University of California, San Diego, has used high-throughput DNA sequencing and machine learning to systematically map the sequence features of “initiator” (Inr) elements in human genes. The study, spearheaded by graduate student Torrey Rhyne-Carrigg, analyzed the transcriptional activity of approximately 500,000 Inr sequence variants, training a highly predictive AI model in the process.
The findings reveal that about 60% of human genes harbor Inr sequences, and the team also identified a novel type of TATA-associated Inr. These insights deepen our understanding of gene expression regulation, help assess the impact of disease-causing mutations, and pave the way for designing synthetic promoters.
Inr elements are short DNA motifs located at transcription start sites, playing a critical role in initiating RNA synthesis. While their importance has been known for decades, the sheer diversity and context-dependent behavior of these sequences made them difficult to characterize on a genome-wide scale. By leveraging AI, the team could move beyond traditional consensus models to capture the full complexity of Inr activity.
“We wanted to move away from approximations and get a complete, quantitative picture of what makes a functional initiator,” said Rhyne-Carrigg. “Machine learning allowed us to see patterns that were previously hidden in the noise.”
The AI model, trained on massive datasets of sequence variants, can now predict with high accuracy whether any given DNA sequence will act as a functional Inr. This capability holds promise for both basic biology and clinical applications. For example, researchers can now evaluate whether mutations near transcription start sites—often found in cancer or developmental disorders—disrupt Inr function, offering new clues about disease mechanisms.
Moreover, the discovery of a TATA-associated Inr variant suggests that some genes rely on a coordinated interplay between two core promoter elements, which could have implications for how cells fine-tune gene expression in different contexts.
The study also demonstrates a broader methodological shift: combining high-throughput experiments with AI-driven analysis is becoming a powerful standard for deciphering regulatory genomics. With the growing availability of sequencing technologies and computational tools, such approaches are likely to become routine in molecular biology.
Looking ahead, the team plans to extend their AI framework to other core promoter elements and to test their models in living organisms. According to Kadonaga, “This is just the beginning. We are now equipped to decode the regulatory language of the human genome at a scale and precision that was unimaginable just a few years ago.”
The work was supported by the National Institutes of Health and was published in the journal Nature Genetics under the title “Machine learning reveals the regulatory grammar of human initiator elements.”

Leave a Reply

Your email address will not be published. Required fields are marked *