Toward Clinically Trustworthy Pathology Foundation Models for Microsatellite Instability Prescreening in Colorectal Cancer
Pathology foundation models have recently emerged as powerful pretrained representations for computational pathology, yet whether complex downstream modeling is still necessary once frozen representations are evaluated under a common downstream framework remains insufficiently understood. We address this question for microsatellite-instability-high (MSI-H) prediction in colorectal cancer by benchmarking nine frozen encoders—five pathology-specific (CONCH, CONCH v1.5, UNI, Virchow2, Phikon) and four conventional vision backbones (ResNet18, ResNet50, ViT-B/16, ConvNeXt-Tiny)—for colorectal cancer histopathology under a patient-level framework, using TCGA-COAD/READ as the development cohort and CPTAC-COAD as an independent external cohort. For each encoder, H&E tiles were embedded without fine-tuning, mean-pooled to slide and patient representations, and evaluated with the same logistic regression linear probe, alongside nonlinear classifiers and a gated-attention multiple-instance learning (MIL) comparator. The complete MANTIS-defined 206-patient cohort (82 MSI-H, 124 MSS) is reported as the primary development benchmark. A label provenance audit identified 15 of 206 development cohort patients (7.3%) with cross-source MSI provenance discordance, and the concordant-label 191-patient cohort (68 MSI-H, 123 MSS) is reported as a secondary sensitivity analysis. In the primary cohort, the selected pathology-specific encoders had a higher mean pooled out-of-fold AUROC than the selected conventional vision backbones (0.827 vs. 0.773; paired-bootstrap difference, +0.054; 95% interval, +0.015 to +0.094), which is reported as a descriptive benchmark rather than a formal inference about the model classes. Virchow2 and UNI produced the strongest linear probe discrimination (AUROC 0.861 and 0.855, respectively), with no consistent gain from the tested nonlinear classifiers or attention–MIL configurations, and representational similarity analysis (linear centered kernel alignment) confirmed that encoders occupy distinct feature geometries rather than converging to a shared representation. In external validation on CPTAC-COAD (105 patients; 24 MSI-H, 81 MSS), selected TCGA-trained models retained variable discrimination, including CONCH attention–MIL AUROC 0.880 and UNI linear probe AUROC 0.859, but external calibration and threshold behavior varied substantially; for CONCH attention–MIL, the development-derived threshold did not transport under the external ensemble implementation (specificity 0.000), whereas threshold-only local adaptation on a small external subset improved median specificity to 0.686 without updating model weights. These findings indicate that modern pathology foundation models can encode MSI-associated morphology in frozen representations under this benchmark, while decision threshold transportability and multimodal or explainability extensions remain open questions for future work.
Authors
- Nguyen Quoc Khanh Le (ORCID: https://orcid.org/0000-0003-4896-7926)
- Kim Ngân Lý
- Nadine Huyen Nguyen
Institutions
- Can Tho University (VN)
- Fulbright University Vietnam (VN)
- Taipei Medical University (TW)
Publication Details
- Journal
- Computers
- Published
- 2026-09-17
- DOI
- https://doi.org/10.3390/computers15090626
- Primary Topic
- AI in cancer detection
- Type
- article
- Field-Weighted Citation Impact
- 0.00
Funders
- Ministry of Science and Technology, Taiwan