Compression-Induced Representation Drift in Pathology Foundation Models
Whole-slide image (WSI) compression is a fundamental requirement in digital pathology. Yet, current validation metrics, such as PSNR and pathologist agreement, were developed before the emergence of pathology foundation models (PFMs). PFMs capture fine-grained tissue details in high-dimensional spaces that can be disrupted by compression artifacts invisible to the human eye, potentially affecting downstream tasks. To the best of our knowledge, we present a systematic study of model- and dataset-specific embedding-drift indicators using representation drift (ΔE), a cosine-based similarity measure, across clinically motivated compression ratios. We evaluate three vision transformer models of identical ViT-Large architecture but different training data: DINOv2 (general vision, 142 million natural images), UNI (pathology, over 100,000 clinical WSIs), and Phikon-v2 (a second PFM). We use 4000 TCGA-BRCA H&E tumor tiles from two WSIs at six compression ratios, ranging from lossless to 80:1. The results demonstrate that UNI first exceeds the predefined ΔE>0.01 sentinel criterion at the tested CR 10:1 operating point (PSNR =37.76 dB, usually considered excellent). DINOv2 first exceeds the same sentinel criterion at CR 20:1. At CR 20:1, UNI’s ΔE=0.109 while DINOv2’s is 0.012, a nearly tenfold difference at the same image quality. At CR 80:1, UNI’s cosine similarity drops to 0.316 while DINOv2 retains 0.895. Phikon-v2 first exceeds the same sentinel criterion at CR 10:1 and tracks UNI far more closely than DINOv2 (ΔE=0.311 versus 0.037 at CR 40:1). The two pathology-pretrained models exhibit substantially greater drift than DINOv2, a pattern consistent with an association between pathology-domain training and increased compression sensitivity. However, checkpoint-specific factors such as training objectives, preprocessing, learned invariances, embedding normalization, and training data may also contribute. Additional tests on a balanced 500-tile TCGA-LUAD pilot set, denser compression ratios, and an alternative JPEG2000 encoder show similar model-dependent drift patterns; however, the empirical drift values and operating criteria should be interpreted as model- and dataset-specific rather than as generalizable thresholds across tissues, institutions, scanners, or staining protocols. We introduce the Rate Distortion Representation (RDR) curve as a model-aware evaluation tool that reveals this representation blind spot: a compression range where PSNR remains conventionally acceptable while embedding geometry is substantially altered. These results are model- and dataset-specific and should not be interpreted as universal compression thresholds or clinical safety boundaries. Overall, the findings support the use of representation-level robustness assessment as a complement to image-level fidelity metrics in AI-oriented digital pathology workflows.
Authors
- Mahmoud R. El-Sakka (ORCID: https://orcid.org/0000-0002-7877-8280)
- Mahmud Hasan (ORCID: https://orcid.org/0000-0003-3338-8163)
- M. Omor Faruk (ORCID: https://orcid.org/0009-0008-8609-3477)
Institutions
- Western University (CA)
- University of Waterloo (CA)
Publication Details
- Journal
- Electronics
- Published
- 2026-09-15
- DOI
- https://doi.org/10.3390/electronics15184186
- Primary Topic
- AI in cancer detection
- Type
- article
- Field-Weighted Citation Impact
- 0.00