meth-SemiCancer2: a cancer subtyping framework leveraging domain adaptation with contrastive learning for batch effect correction in DNA methylation profiles
Cancer subtyping is essential for precision oncology, enabling early diagnosis, targeted therapy, and improved prognosis. DNA methylation is a powerful molecular marker for subtype classification, but its application is hindered by high dimensionality, batch effects, and the scarcity of subtype-labeled datasets. We previously developed meth-SemiCancer, a semi-supervised framework that leverages unlabeled methylation data to improve subtype prediction. While effective, meth-SemiCancer did not explicitly address batch effect correction or robust representation learning, limiting its generalizability across cohorts. Without addressing these challenges, technical noise can obscure true biological signals, leading to spurious subtype boundaries and limiting reproducibility across studies. To overcome these limitations, we developed meth-SemiCancer2, an enhanced framework that employs domain adaptation coupled with contrastive learning, followed by semi-supervised fine-tuning with subtype alignment. Specifically, meth-SemiCancer2 first pretrains the model on labeled source datasets to capture subtype-specific patterns, then applies adversarial training to align distributions between source and unlabeled target cohorts. Contrastive learning is subsequently introduced to strengthen domain-invariant feature extraction, and finally, semi-supervised fine-tuning with subtype alignment iteratively refines pseudo-labels to improve subtype separation across cohorts. Through this multi-phase, the model corrects batch effects while learning robust and discriminative subtype representations from high-dimensional methylation profiles. Evaluations across seven cancer types-breast cancer, colon adenocarcinoma, prostate adenocarcinoma, glioblastoma, renal cell carcinoma, thyroid carcinoma, and low-grade glioma-demonstrated that meth-SemiCancer2 consistently outperformed recently proposed deep learning-based DNA methylation subtype classification frameworks, including our previous meth-SemiCancer model, a representative domain adaptation method, two state-of-the-art DNA methylation batch correction approaches, and conventional machine learning classifiers. In labeled source datasets, 10-fold cross-validation verified its superior accuracy and reliability in subtype prediction. When applied to target cohorts, the model generalized effectively, producing accurate subtype assignments with clearly delineated clusters. Our evaluations further indicated that meth-SemiCancer2 sustained stable performance even when only fractions of the source or target data were available, highlighting its utility under data-limited conditions. meth-SemiCancer2 is publicly accessible at https://github.com/cbi-bioinfo/meth-semicancer2 .
Authors
- Heejoon Chae (ORCID: https://orcid.org/0000-0002-0960-5829)
- Joung Min Choi
Institutions
- Sookmyung Women's University (KR)
- Virginia Tech (US)
Publication Details
- Journal
- BMC Bioinformatics
- Published
- 2026-09-28
- DOI
- https://doi.org/10.1186/s12859-026-06656-0
- Primary Topic
- Epigenetics and DNA Methylation
- Type
- article
- Field-Weighted Citation Impact
- 0.00