Transcriptome-Based Subtype Discovery in Breast Cancer Using Variational Representation Learning
Breast cancer is a heterogeneous disease whose molecular complexity complicates therapeutic stratification. Although PAM50 is a widely used molecular subtyping framework, discrete subtype assignments may underrepresent transitional molecular states and within-subtype heterogeneity relevant to treatment. This study develops a biologically interpretable representation learning framework that projects bulk RNA-seq profiles into nonlinear latent spaces and applies unsupervised hierarchical clustering, cross-validated against k-means and Gaussian mixture models (GMMs), to investigate additional tumor structure. The pipeline integrates a variational autoencoder (VAE), t-SNE/UMAP visualization, and functional (differential expression, pathway enrichment) and immune (ESTIMATE, quanTIseq) characterization of the resulting clusters. From a source pool of 19,131 transcriptomic profiles, a primary analytical cohort of 8744 unique, non-duplicated, PAM50-labeled tumor and normal samples was used for VAE training and all downstream analyses, without class balancing or oversampling. Across this primary cohort, tissue of origin and malignant-versus-normal status—not PAM50 subtype—consistently dominated clustering structure across every representation (PCA−40; VAEZ=10,20,40) and every clustering algorithm tested, a pattern reproduced even in the raw gene expression space. Isolating the resulting 835-sample malignant subpopulation and reclustering it in isolation revealed subtype-specific structure obscured at the whole-cohort level: Basal-like disease separated cleanly under every representation, algorithm, and gene-panel size tested (83–98% purity); Luminal A resolved into a distinct cluster plus a continuum bridging Luminal B; HER2-enriched concentrated but only inconsistently achieved cluster-majority status; and Luminal B never achieved outright cluster majority under any configuration tested, behaving as a shared boundary population rather than an independently separable subtype. Among malignant-subset representations, VAEZ=10 achieved the strongest internal validation metrics, the highest concordance with PAM50 (AdjustedRandIndex=0.377), and—together with VAEZ=40—a concordance index exceeding the PAM50-only baseline (0.619 and 0.643 versus 0.606). Immune deconvolution showed that immune-hot and immune-cold status cut across, rather than align with, PAM50 boundaries, including a regulatory T-cell/M1-macrophage-skewed HER2-enriched/Luminal B cluster and a CD8+-poor, B-cell-skewed Basal-like cluster. Bootstrap resampling and gene-panel-size sensitivity analyses confirmed that cluster stability tracked PAM50 purity, with Basal-like separation robust across all conditions and the HER2/Luminal B boundary consistently the least stable. These internally derived findings indicate that nonlinear representation learning can recover biologically coherent structure that refines, rather than replaces, PAM50 classification, though independent-cohort and spatial or single-cell validation are required before clinical interpretation.
Authors
- Hazem Hiary (ORCID: https://orcid.org/0000-0002-0306-5294)
- Loai Alnemer (ORCID: https://orcid.org/0000-0002-1208-9861)
- Al Hanouf Al Khawaldeh
Institutions
- University of Jordan (JO)
Publication Details
- Journal
- BioMedInformatics
- Published
- 2026-10-04
- DOI
- https://doi.org/10.3390/biomedinformatics6050086
- Primary Topic
- Breast Cancer Treatment Studies
- Type
- article
- Field-Weighted Citation Impact
- 0.00