EnhancerDetector: enhancer discovery from human to fly via interpretable deep learning

Deciphering how enhancers encode regulatory information in DNA remains a central genomics challenge, as sequencing outpaces functional annotation. A key question is whether enhancers possess recurring, sequence-based enhancer-associated features—here termed “enhancerness”—that distinguish them from other genomic regions across species, cell types, and experimental assays.Confirming its existence and learnability is both biologically fundamental and essential for scalable genome annotation. We introduce EnhancerDetector , a convolutional neural network-based framework for cross-species enhancer prediction that combines high accuracy with biological interpretability. Trained on human data, EnhancerDetector achieves strong performance across human, mouse, and fly datasets, consistently outperforming existing methods in precision and F1. It generalizes to datasets generated using diverse experimental assays. Unlike chromatin feature-based predictors requiring complex post hoc thresholding, EnhancerDetector directly outputs enhancer probability scores from short sequence windows, simplifying enhancer discovery workflows. An ensemble strategy further improves prediction reliability by reducing false positives. EnhancerDetector supports fine-tuning on new species and retains strong performance even when adapted with as few as 20,000 enhancer sequences, making it ideal for newly sequenced genomes with limited experimental data. For interpretability and visualization, we apply class activation maps to identify sequence regions predictive of enhancer activity. Experimental validation in transgenic flies confirms the predictive power of EnhancerDetector : five of six tested candidates drove reporter expression, and four exhibited expression patterns supported by prior literature. These analyses identify distinct sequence and contextual features associated with enhancer activity, collectively referred to here as “enhancerness,” suggesting that enhancers share learnable sequence characteristics. EnhancerDetector provides a sequence-based and interpretable framework for enhancer discovery across species and experimental contexts. Our results support the presence of recurring enhancer-associated sequence features that can be learned from DNA sequence and transferred across genomes. By combining cross-species prediction, fine-tuning, model interpretation, and experimental validation, EnhancerDetector offers a practical approach for prioritizing candidate enhancers in well-studied genomes. These features position EnhancerDetector as a useful first-stage annotation tool for identifying putative enhancers in newly sequenced genomes with limited experimental data.

Authors

Institutions

Publication Details

Journal
BMC Genomics
Published
2026-09-01
DOI
https://doi.org/10.1186/s12864-026-13262-0
Primary Topic
Zebrafish Biomedical Research Applications
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

EnhancerDetector: enhancer discovery from human to fly via interpretable deep learning

Hani Z. Girgis, Marc S. Halfon, Geyenna Sterling-Lentsch, Luis M. Solis
BMC Genomics
Zebrafish Biomedical Research Applications
article

EnhancerDetector: enhancer discovery from human to fly via interpretable deep learning

Hani Z. Girgis, Marc S. Halfon, Geyenna Sterling-Lentsch, Luis M. Solis
article en

Abstract

Deciphering how enhancers encode regulatory information in DNA remains a central genomics challenge, as sequencing outpaces functional annotation. A key question is whether enhancers possess recurring, sequence-based enhancer-associated features—here termed “enhancerness”—that distinguish them from other genomic regions across species, cell types, and experimental assays.Confirming its existence and learnability is both biologically fundamental and essential for scalable genome annotation. We introduce EnhancerDetector , a convolutional neural network-based framework for cross-species enhancer prediction that combines high accuracy with biological interpretability. Trained on human data, EnhancerDetector achieves strong performance across human, mouse, and fly datasets, consistently outperforming existing methods in precision and F1. It generalizes to datasets generated using diverse experimental assays. Unlike chromatin feature-based predictors requiring complex post hoc thresholding, EnhancerDetector directly outputs enhancer probability scores from short sequence windows, simplifying enhancer discovery workflows. An ensemble strategy further improves prediction reliability by reducing false positives. EnhancerDetector supports fine-tuning on new species and retains strong performance even when adapted with as few as 20,000 enhancer sequences, making it ideal for newly sequenced genomes with limited experimental data. For interpretability and visualization, we apply class activation maps to identify sequence regions predictive of enhancer activity. Experimental validation in transgenic flies confirms the predictive power of EnhancerDetector : five of six tested candidates drove reporter expression, and four exhibited expression patterns supported by prior literature. These analyses identify distinct sequence and contextual features associated with enhancer activity, collectively referred to here as “enhancerness,” suggesting that enhancers share learnable sequence characteristics. EnhancerDetector provides a sequence-based and interpretable framework for enhancer discovery across species and experimental contexts. Our results support the presence of recurring enhancer-associated sequence features that can be learned from DNA sequence and transferred across genomes. By combining cross-species prediction, fine-tuning, model interpretation, and experimental validation, EnhancerDetector offers a practical approach for prioritizing candidate enhancers in well-studied genomes. These features position EnhancerDetector as a useful first-stage annotation tool for identifying putative enhancers in newly sequenced genomes with limited experimental data.

BMC Genomics
Texas A&M University – Kingsville (US), University at Buffalo, State University of New York (US)
Openalex Percentile: Top 14%
Zebrafish Biomedical Research Applications
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.