Decoding promoter activity from DNA sequence using pre-trained language models

C Christophe Jung

Abstract

Abstract Promoter architecture plays a central role in transcriptional regulation, but predicting promoter activity directly from DNA sequence remains challenging. Here, we tested whether transformer-based DNA language models can learn regulatory logic encoded in Drosophila core promoters. We fine-tuned the pretrained DNA language model DNABERT-2 using a synthetic core promoter dataset measured in S2 cells with luciferase reporter assays. The model predicted promoter activity well when biological replicates were split between training and test data (R² ≈ 0.91), and retained meaningful performance when test promoter sequences were fully excluded from training (R² ≈ 0.64). Model interpretation using SHapley Additive exPlanations (SHAP) 1 showed that predictive sequence features matched known promoter elements, including INR, TATA box, DRE, Ohler and MTE/DPE motifs, with position-dependent effects consistent with promoter architecture. Incorporating hormonal activation and nucleosomal context enabled sequence and biological context to be modeled in a unified framework. Gene-wise cross-validation showed promoter-specific generalization across most promoters with promoter-specific differences in accuracy. Applied without retraining to independent Drosophila embryo promoter data, the model captured partial in vivo activity trends. These results show that DNA language models can learn interpretable promoter sequence rules from controlled datasets, while accurate in vivo prediction will require broader regulatory context.

Article Details

Volume / Issue Vol. 16, Issue 1
Published July 14, 2026
ISSN 2045-2322
Publisher Nature Portfolio

Journal Info

Scientific Reports

Nature Portfolio

ISSN: 2045-2322 Open Access Life Sciences

Authors (1)

C

Christophe Jung