PathScaleBench: A Multidimensional Benchmark of Cross-Scale Representation Behavior and Magnification-Shift Robustness in Pathology Foundation Models
Abstract Pathology foundation models (PFMs) are widely used as frozen encoders, but their behavior under changes in histologic magnification remains poorly characterized. We developed PathScaleBench to quantify cross-scale representation behavior and determine whether it is associated with downstream magnification-shift robustness. Seventy-one audited native-40 × The Cancer Genome Atlas whole-slide images from eight cancer types were sampled at the same tissue centers at 40 × , 10 × , and 2.5 × using a native-field-of-view (native-FOV) protocol. Eight PFMs were evaluated using centered kernel alignment (CKA), relative retention, bidirectional cross-scale retrieval, cancer retrieval, and neighborhood preservation. Eight-class linear probes were trained at one magnification and tested at all three using ten repetitions of five-fold patient-grouped cross-validation. Model rankings varied by metric. Phikon achieved the highest mean CKA (0.714), relative retention (0.774), and neighborhood preservation (0.248); Virchow2 achieved the highest cross-scale retrieval similarity (0.643); and Midnight achieved the highest cancer Recall@1 (0.845). No representation metric was significantly associated with Recall@1 at the model level after multiplicity correction. Midnight showed the lowest magnification generalization gap (0.082), followed by Virchow2 (0.181) and UNI2 (0.224). Across 1704 case-model-scale-pair observations, pair-matched retrieval similarity was associated with corresponding shift robustness (Spearman ρ = 0.159, 95% confidence interval 0.059–0.264; Holm-adjusted p = 0.0021). Cross-scale PFM behavior is multidimensional; global alignment, local correspondence, biological identity, neighborhood topology, and downstream robustness are not interchangeable. No model dominated all dimensions. Representation profiles can support model selection, but magnification robustness should be tested directly.
Authors
- Fei Su (ORCID: https://orcid.org/0000-0002-6850-9773)
Publication Details
- Journal
- Journal of Imaging Informatics in Medicine
- Published
- 2026-10-07
- DOI
- https://doi.org/10.1007/s10278-026-02354-8
- Primary Topic
- AI in cancer detection
- Type
- article
- Field-Weighted Citation Impact
- 0.00