GP-DHT: A dual-head transformer with supervised contrastive learning for shared-gene pretraining and single-cell GRN prediction
Gene Regulatory Networks (GRNs) are crucial for understanding cell fate decisions and disease mechanisms. However, inferring GRNs from single-cell RNA sequencing remains challenging due to noise, data sparsity, and species-specific distribution variations. We propose GP-DHT (Gene Pair–Dual Head Transformer), a single-cell GRN inference framework that pre-trains gene representations on a unified shared gene space using the BEELINE datasets of humans and mice. It models gene-cell expression relationships through a multi-relational heterograph and predicts transcription factor (TF)-target regulatory links using a dual-head Transformer. These two heads correspond to complementary task objectives: a classification head for regulatory link prediction and a projection head for gene pair-level supervised contrastive learning. This design enables the model to optimize both prediction accuracy and representation separability in a sparse single-cell setting. The human-mouse shared-gene stage is used to initialize expression-derived gene representations in a harmonized gene space. During dataset-wise supervised training, GenePairSupCon is applied within each training fold to regularize gene-pair representations, followed by fine-tuning and evaluation within each BEELINE dataset. To enhance reproducibility and interpretability, we further clarify the mapping of orthologous genes and gene symbol unification, TF-level data partitioning and leakage control, negative label sensitivity, and quantitative embedding quality assessment. Hard negative example experiments show that negative label construction has a substantial impact on task difficulty, while supervised contrastive learning significantly improves the separability of embeddings, as measured by Silhouette score, Davies-Bouldin index, Calinski-Harabasz score, and inter/intra distance ratio. GP-DHT can also recover prior-supported candidate regulatory modules and provide additional interpretable support through post-hoc pair-level attribution analysis on frozen predictors. In summary, these results support GP-DHT as a robust shared gene representation learning framework for dataset-level single-cell GRN prediction.
Authors
- Xiang Cheng (ORCID: https://orcid.org/0009-0006-1919-3414)
- Qingzhi Yu
- Shuai Yan
- Wenfeng Dai
Institutions
- Jingdezhen Ceramic Institute (CN)
Publication Details
- Journal
- PLoS ONE
- Published
- 2026-10-06
- DOI
- https://doi.org/10.1371/journal.pone.0359164
- Primary Topic
- Gene Regulatory Network Analysis
- Type
- article
- Field-Weighted Citation Impact
- 0.00