Calibration-Aware Patient-Level Prediction of Source-Defined MSIMUT Status in Colorectal Cancer from Whole-Slide Histopathology Images Using Foundation-Model Embeddings

Background/Objectives: Microsatellite instability (MSI) is clinically important in colorectal cancer, but definitive assessment requires molecular or immunohistochemical testing. This study developed a leakage-safe, calibration-aware patient-level framework for predicting a source-defined MSIMUT category from hematoxylin-and-eosin whole-slide-derived patches using fixed Phikon foundation-model embeddings. Methods: The TCGA-derived colorectal cohort comprised 360 patients, including 65 source-defined positive and 295 MSS cases, represented by 192,312 pre-extracted patches. Patch embeddings were aggregated using mean, max, and mean-plus-max pooling, and downstream classifiers were evaluated under patient-level cross-validation with calibration, threshold, decision-curve, clinicopathological, source-provenance, and patch-count sensitivity analyses. Results: In the primary random patient-level evaluation, logistic regression with mean-plus-max pooling achieved a mean ROC-AUC of 0.915 and PR-AUC of 0.809. Fully nested Platt calibration yielded pooled out-of-sample ROC-AUC 0.913, PR-AUC 0.796, Brier score 0.073, calibration slope 1.037, and calibration-in-the-large 0.023. Within the internally cross-validated mean-plus-max framework, an exploratory sensitivity ≥ 0.95 operating point yielded sensitivity 0.969, specificity 0.393, and negative predictive value 0.983. Provenance reconstruction identified source-partition-dependent patch-count structure. In a source-TRAIN to source-TEST holdout, mean-plus-max ROC-AUC decreased to 0.788, while patch count alone was near chance in source-TEST (ROC-AUC 0.517) and explicit patch-count addition did not improve discrimination. Conclusions: These findings support the methodological feasibility of fixed foundation-model embeddings for internal MSI-related pre-screening research, while showing that apparent performance is sensitive to source structure. Independent multicenter validation with clinically adjudicated MSI-H/dMMR endpoints remains necessary before clinical application.

Authors

Institutions

Publication Details

Journal
Diagnostics
Published
2026-10-09
DOI
https://doi.org/10.3390/diagnostics16203271
Primary Topic
AI in cancer detection
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Calibration-Aware Patient-Level Prediction of Source-Defined MSIMUT Status in Colorectal Cancer from Whole-Slide Histopathology Images Using Foundation-Model Embeddings

Deniz Özkan Vardar, Necati Vardar
Diagnostics
AI in cancer detection
article

Calibration-Aware Patient-Level Prediction of Source-Defined MSIMUT Status in Colorectal Cancer from Whole-Slide Histopathology Images Using Foundation-Model Embeddings

Deniz Özkan Vardar, Necati Vardar
article en

Abstract

Background/Objectives: Microsatellite instability (MSI) is clinically important in colorectal cancer, but definitive assessment requires molecular or immunohistochemical testing. This study developed a leakage-safe, calibration-aware patient-level framework for predicting a source-defined MSIMUT category from hematoxylin-and-eosin whole-slide-derived patches using fixed Phikon foundation-model embeddings. Methods: The TCGA-derived colorectal cohort comprised 360 patients, including 65 source-defined positive and 295 MSS cases, represented by 192,312 pre-extracted patches. Patch embeddings were aggregated using mean, max, and mean-plus-max pooling, and downstream classifiers were evaluated under patient-level cross-validation with calibration, threshold, decision-curve, clinicopathological, source-provenance, and patch-count sensitivity analyses. Results: In the primary random patient-level evaluation, logistic regression with mean-plus-max pooling achieved a mean ROC-AUC of 0.915 and PR-AUC of 0.809. Fully nested Platt calibration yielded pooled out-of-sample ROC-AUC 0.913, PR-AUC 0.796, Brier score 0.073, calibration slope 1.037, and calibration-in-the-large 0.023. Within the internally cross-validated mean-plus-max framework, an exploratory sensitivity ≥ 0.95 operating point yielded sensitivity 0.969, specificity 0.393, and negative predictive value 0.983. Provenance reconstruction identified source-partition-dependent patch-count structure. In a source-TRAIN to source-TEST holdout, mean-plus-max ROC-AUC decreased to 0.788, while patch count alone was near chance in source-TEST (ROC-AUC 0.517) and explicit patch-count addition did not improve discrimination. Conclusions: These findings support the methodological feasibility of fixed foundation-model embeddings for internal MSI-related pre-screening research, while showing that apparent performance is sensitive to source structure. Independent multicenter validation with clinically adjudicated MSI-H/dMMR endpoints remains necessary before clinical application.

DiagnosticsVol. 16(20)
Lokman Hekim Üniversitesi (TR), KTO Karatay University (TR)
Openalex Percentile: Top 13%
AI in cancer detection
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.