Transcriptome-Based Subtype Discovery in Breast Cancer Using Variational Representation Learning

Breast cancer is a heterogeneous disease whose molecular complexity complicates therapeutic stratification. Although PAM50 is a widely used molecular subtyping framework, discrete subtype assignments may underrepresent transitional molecular states and within-subtype heterogeneity relevant to treatment. This study develops a biologically interpretable representation learning framework that projects bulk RNA-seq profiles into nonlinear latent spaces and applies unsupervised hierarchical clustering, cross-validated against k-means and Gaussian mixture models (GMMs), to investigate additional tumor structure. The pipeline integrates a variational autoencoder (VAE), t-SNE/UMAP visualization, and functional (differential expression, pathway enrichment) and immune (ESTIMATE, quanTIseq) characterization of the resulting clusters. From a source pool of 19,131 transcriptomic profiles, a primary analytical cohort of 8744 unique, non-duplicated, PAM50-labeled tumor and normal samples was used for VAE training and all downstream analyses, without class balancing or oversampling. Across this primary cohort, tissue of origin and malignant-versus-normal status—not PAM50 subtype—consistently dominated clustering structure across every representation (PCA−40; VAEZ=10,20,40) and every clustering algorithm tested, a pattern reproduced even in the raw gene expression space. Isolating the resulting 835-sample malignant subpopulation and reclustering it in isolation revealed subtype-specific structure obscured at the whole-cohort level: Basal-like disease separated cleanly under every representation, algorithm, and gene-panel size tested (83–98% purity); Luminal A resolved into a distinct cluster plus a continuum bridging Luminal B; HER2-enriched concentrated but only inconsistently achieved cluster-majority status; and Luminal B never achieved outright cluster majority under any configuration tested, behaving as a shared boundary population rather than an independently separable subtype. Among malignant-subset representations, VAEZ=10 achieved the strongest internal validation metrics, the highest concordance with PAM50 (AdjustedRandIndex=0.377), and—together with VAEZ=40—a concordance index exceeding the PAM50-only baseline (0.619 and 0.643 versus 0.606). Immune deconvolution showed that immune-hot and immune-cold status cut across, rather than align with, PAM50 boundaries, including a regulatory T-cell/M1-macrophage-skewed HER2-enriched/Luminal B cluster and a CD8+-poor, B-cell-skewed Basal-like cluster. Bootstrap resampling and gene-panel-size sensitivity analyses confirmed that cluster stability tracked PAM50 purity, with Basal-like separation robust across all conditions and the HER2/Luminal B boundary consistently the least stable. These internally derived findings indicate that nonlinear representation learning can recover biologically coherent structure that refines, rather than replaces, PAM50 classification, though independent-cohort and spatial or single-cell validation are required before clinical interpretation.

Authors

Institutions

Publication Details

Journal
BioMedInformatics
Published
2026-10-04
DOI
https://doi.org/10.3390/biomedinformatics6050086
Primary Topic
Breast Cancer Treatment Studies
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Transcriptome-Based Subtype Discovery in Breast Cancer Using Variational Representation Learning

Hazem Hiary, Loai Alnemer, Al Hanouf Al Khawaldeh
BioMedInformatics
Breast Cancer Treatment Studies
article

Transcriptome-Based Subtype Discovery in Breast Cancer Using Variational Representation Learning

Hazem Hiary, Loai Alnemer, Al Hanouf Al Khawaldeh
article en

Abstract

Breast cancer is a heterogeneous disease whose molecular complexity complicates therapeutic stratification. Although PAM50 is a widely used molecular subtyping framework, discrete subtype assignments may underrepresent transitional molecular states and within-subtype heterogeneity relevant to treatment. This study develops a biologically interpretable representation learning framework that projects bulk RNA-seq profiles into nonlinear latent spaces and applies unsupervised hierarchical clustering, cross-validated against k-means and Gaussian mixture models (GMMs), to investigate additional tumor structure. The pipeline integrates a variational autoencoder (VAE), t-SNE/UMAP visualization, and functional (differential expression, pathway enrichment) and immune (ESTIMATE, quanTIseq) characterization of the resulting clusters. From a source pool of 19,131 transcriptomic profiles, a primary analytical cohort of 8744 unique, non-duplicated, PAM50-labeled tumor and normal samples was used for VAE training and all downstream analyses, without class balancing or oversampling. Across this primary cohort, tissue of origin and malignant-versus-normal status—not PAM50 subtype—consistently dominated clustering structure across every representation (PCA−40; VAEZ=10,20,40) and every clustering algorithm tested, a pattern reproduced even in the raw gene expression space. Isolating the resulting 835-sample malignant subpopulation and reclustering it in isolation revealed subtype-specific structure obscured at the whole-cohort level: Basal-like disease separated cleanly under every representation, algorithm, and gene-panel size tested (83–98% purity); Luminal A resolved into a distinct cluster plus a continuum bridging Luminal B; HER2-enriched concentrated but only inconsistently achieved cluster-majority status; and Luminal B never achieved outright cluster majority under any configuration tested, behaving as a shared boundary population rather than an independently separable subtype. Among malignant-subset representations, VAEZ=10 achieved the strongest internal validation metrics, the highest concordance with PAM50 (AdjustedRandIndex=0.377), and—together with VAEZ=40—a concordance index exceeding the PAM50-only baseline (0.619 and 0.643 versus 0.606). Immune deconvolution showed that immune-hot and immune-cold status cut across, rather than align with, PAM50 boundaries, including a regulatory T-cell/M1-macrophage-skewed HER2-enriched/Luminal B cluster and a CD8+-poor, B-cell-skewed Basal-like cluster. Bootstrap resampling and gene-panel-size sensitivity analyses confirmed that cluster stability tracked PAM50 purity, with Basal-like separation robust across all conditions and the HER2/Luminal B boundary consistently the least stable. These internally derived findings indicate that nonlinear representation learning can recover biologically coherent structure that refines, rather than replaces, PAM50 classification, though independent-cohort and spatial or single-cell validation are required before clinical interpretation.

BioMedInformaticsVol. 6(5)
University of Jordan (JO)
Openalex Percentile: Top 17%
Breast Cancer Treatment Studies
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.