AI-Based Prognostic Risk Stratification of Adult-Type Diffuse Glioma Patients from H&E-Stained Slides

Background/Objectives: Adult-type diffuse gliomas (ADGs) are aggressive brain tumors with highly variable clinical outcomes. Prognostic assessment directly from routine hematoxylin and eosin (H&E)-stained whole-slide images (WSIs) could support personalized decision-making. Methods: To our knowledge, this is the first systematic benchmarking of pathology foundation models (FMs) and multiple instance learning (MIL) methods for overall survival (OS) risk stratification in ADGs under the 2021 World Health Organization (WHO) classification. A retrospective, multi-institutional cohort of 817 subjects from The Cancer Genome Atlas (TCGA) was analyzed. Following tissue segmentation and patch-level curation, we evaluated 77 pathology FM-MIL combinations (11 FMs × 7 MIL methods), together with seven ResNet-50 baseline configurations (84 encoder-MIL configurations in total), using patient-level 10-fold cross-validation. A perturbation-based evaluation assessed attention mechanism faithfulness. Results: UNI2 paired with MambaMIL achieved the highest cross-validated concordance index (C-index = 0.77), whereas its time-dependent AUC was modest (0.60–0.61 at 12–36 months), and several combinations performed comparably. Kaplan–Meier analysis demonstrated significant separation between high- and low-risk groups (log-rank p = 9.16 × 10−17; hazard ratio [HR] = 2.90, 95% confidence interval [CI]: 2.23–3.76), with a median OS of 16.6 months (95% CI: 15.0–20.4) versus 63.5 months (95% CI: 43.9-NR), respectively. Tertile stratification confirmed a monotonic OS gradient from 94.5 months to 14.7 months (p = 3.25 × 10−17). Risk stratification within individual WHO 2021 types was not significant. Most domain-specific FMs outperformed the ImageNet-pretrained baseline, and the removal of highly attended patches reduced performance, supporting attention faithfulness. Conclusions: The AI-derived risk score remained associated with OS after adjustment for individual clinical and molecular variables, but independent prognostic value was not demonstrated in the fully adjusted model (HR = 1.32, 95% CI: 0.94–1.85, p = 0.104), and its discrimination largely reflects differences between WHO 2021 types. This benchmarking framework provides guidance for FM-MIL design in computational pathology applications targeting survival prediction; external validation is required before clinical application.

Authors

Institutions

Publication Details

Journal
Cancers
Published
2026-10-09
DOI
https://doi.org/10.3390/cancers18203259
Primary Topic
AI in cancer detection
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

AI-Based Prognostic Risk Stratification of Adult-Type Diffuse Glioma Patients from H&E-Stained Slides

Shubham Innani, Spyridon Bakas, Carla Pitarch, Dimitrios Makris et al.
Cancers
AI in cancer detection
article

AI-Based Prognostic Risk Stratification of Adult-Type Diffuse Glioma Patients from H&E-Stained Slides

Shubham Innani, Spyridon Bakas, Carla Pitarch, Dimitrios Makris, Hannah J. Harmsen, Marwan M. Majeed, W. Robert Bell
article en

Abstract

Background/Objectives: Adult-type diffuse gliomas (ADGs) are aggressive brain tumors with highly variable clinical outcomes. Prognostic assessment directly from routine hematoxylin and eosin (H&E)-stained whole-slide images (WSIs) could support personalized decision-making. Methods: To our knowledge, this is the first systematic benchmarking of pathology foundation models (FMs) and multiple instance learning (MIL) methods for overall survival (OS) risk stratification in ADGs under the 2021 World Health Organization (WHO) classification. A retrospective, multi-institutional cohort of 817 subjects from The Cancer Genome Atlas (TCGA) was analyzed. Following tissue segmentation and patch-level curation, we evaluated 77 pathology FM-MIL combinations (11 FMs × 7 MIL methods), together with seven ResNet-50 baseline configurations (84 encoder-MIL configurations in total), using patient-level 10-fold cross-validation. A perturbation-based evaluation assessed attention mechanism faithfulness. Results: UNI2 paired with MambaMIL achieved the highest cross-validated concordance index (C-index = 0.77), whereas its time-dependent AUC was modest (0.60–0.61 at 12–36 months), and several combinations performed comparably. Kaplan–Meier analysis demonstrated significant separation between high- and low-risk groups (log-rank p = 9.16 × 10−17; hazard ratio [HR] = 2.90, 95% confidence interval [CI]: 2.23–3.76), with a median OS of 16.6 months (95% CI: 15.0–20.4) versus 63.5 months (95% CI: 43.9-NR), respectively. Tertile stratification confirmed a monotonic OS gradient from 94.5 months to 14.7 months (p = 3.25 × 10−17). Risk stratification within individual WHO 2021 types was not significant. Most domain-specific FMs outperformed the ImageNet-pretrained baseline, and the removal of highly attended patches reduced performance, supporting attention faithfulness. Conclusions: The AI-derived risk score remained associated with OS after adjustment for individual clinical and molecular variables, but independent prognostic value was not demonstrated in the fully adjusted model (HR = 1.32, 95% CI: 0.94–1.85, p = 0.104), and its discrimination largely reflects differences between WHO 2021 types. This benchmarking framework provides guidance for FM-MIL design in computational pathology applications targeting survival prediction; external validation is required before clinical application.

CancersVol. 18(20)
Kingston University London (GB), Indiana University Indianapolis (US), Indiana University Melvin and Bren Simon Comprehensive Cancer Center, Indiana University School of Medicine, Indiana University – Purdue University Indianapolis (US)
Openalex Percentile: Top 12%
AI in cancer detection
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.