Compression-Induced Representation Drift in Pathology Foundation Models

Whole-slide image (WSI) compression is a fundamental requirement in digital pathology. Yet, current validation metrics, such as PSNR and pathologist agreement, were developed before the emergence of pathology foundation models (PFMs). PFMs capture fine-grained tissue details in high-dimensional spaces that can be disrupted by compression artifacts invisible to the human eye, potentially affecting downstream tasks. To the best of our knowledge, we present a systematic study of model- and dataset-specific embedding-drift indicators using representation drift (ΔE), a cosine-based similarity measure, across clinically motivated compression ratios. We evaluate three vision transformer models of identical ViT-Large architecture but different training data: DINOv2 (general vision, 142 million natural images), UNI (pathology, over 100,000 clinical WSIs), and Phikon-v2 (a second PFM). We use 4000 TCGA-BRCA H&E tumor tiles from two WSIs at six compression ratios, ranging from lossless to 80:1. The results demonstrate that UNI first exceeds the predefined ΔE>0.01 sentinel criterion at the tested CR 10:1 operating point (PSNR =37.76 dB, usually considered excellent). DINOv2 first exceeds the same sentinel criterion at CR 20:1. At CR 20:1, UNI’s ΔE=0.109 while DINOv2’s is 0.012, a nearly tenfold difference at the same image quality. At CR 80:1, UNI’s cosine similarity drops to 0.316 while DINOv2 retains 0.895. Phikon-v2 first exceeds the same sentinel criterion at CR 10:1 and tracks UNI far more closely than DINOv2 (ΔE=0.311 versus 0.037 at CR 40:1). The two pathology-pretrained models exhibit substantially greater drift than DINOv2, a pattern consistent with an association between pathology-domain training and increased compression sensitivity. However, checkpoint-specific factors such as training objectives, preprocessing, learned invariances, embedding normalization, and training data may also contribute. Additional tests on a balanced 500-tile TCGA-LUAD pilot set, denser compression ratios, and an alternative JPEG2000 encoder show similar model-dependent drift patterns; however, the empirical drift values and operating criteria should be interpreted as model- and dataset-specific rather than as generalizable thresholds across tissues, institutions, scanners, or staining protocols. We introduce the Rate Distortion Representation (RDR) curve as a model-aware evaluation tool that reveals this representation blind spot: a compression range where PSNR remains conventionally acceptable while embedding geometry is substantially altered. These results are model- and dataset-specific and should not be interpreted as universal compression thresholds or clinical safety boundaries. Overall, the findings support the use of representation-level robustness assessment as a complement to image-level fidelity metrics in AI-oriented digital pathology workflows.

Authors

Institutions

Publication Details

Journal
Electronics
Published
2026-09-15
DOI
https://doi.org/10.3390/electronics15184186
Primary Topic
AI in cancer detection
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Compression-Induced Representation Drift in Pathology Foundation Models

Mahmoud R. El-Sakka, Mahmud Hasan, M. Omor Faruk
Electronics
AI in cancer detection
article

Compression-Induced Representation Drift in Pathology Foundation Models

Mahmoud R. El-Sakka, Mahmud Hasan, M. Omor Faruk
article en

Abstract

Whole-slide image (WSI) compression is a fundamental requirement in digital pathology. Yet, current validation metrics, such as PSNR and pathologist agreement, were developed before the emergence of pathology foundation models (PFMs). PFMs capture fine-grained tissue details in high-dimensional spaces that can be disrupted by compression artifacts invisible to the human eye, potentially affecting downstream tasks. To the best of our knowledge, we present a systematic study of model- and dataset-specific embedding-drift indicators using representation drift (ΔE), a cosine-based similarity measure, across clinically motivated compression ratios. We evaluate three vision transformer models of identical ViT-Large architecture but different training data: DINOv2 (general vision, 142 million natural images), UNI (pathology, over 100,000 clinical WSIs), and Phikon-v2 (a second PFM). We use 4000 TCGA-BRCA H&E tumor tiles from two WSIs at six compression ratios, ranging from lossless to 80:1. The results demonstrate that UNI first exceeds the predefined ΔE>0.01 sentinel criterion at the tested CR 10:1 operating point (PSNR =37.76 dB, usually considered excellent). DINOv2 first exceeds the same sentinel criterion at CR 20:1. At CR 20:1, UNI’s ΔE=0.109 while DINOv2’s is 0.012, a nearly tenfold difference at the same image quality. At CR 80:1, UNI’s cosine similarity drops to 0.316 while DINOv2 retains 0.895. Phikon-v2 first exceeds the same sentinel criterion at CR 10:1 and tracks UNI far more closely than DINOv2 (ΔE=0.311 versus 0.037 at CR 40:1). The two pathology-pretrained models exhibit substantially greater drift than DINOv2, a pattern consistent with an association between pathology-domain training and increased compression sensitivity. However, checkpoint-specific factors such as training objectives, preprocessing, learned invariances, embedding normalization, and training data may also contribute. Additional tests on a balanced 500-tile TCGA-LUAD pilot set, denser compression ratios, and an alternative JPEG2000 encoder show similar model-dependent drift patterns; however, the empirical drift values and operating criteria should be interpreted as model- and dataset-specific rather than as generalizable thresholds across tissues, institutions, scanners, or staining protocols. We introduce the Rate Distortion Representation (RDR) curve as a model-aware evaluation tool that reveals this representation blind spot: a compression range where PSNR remains conventionally acceptable while embedding geometry is substantially altered. These results are model- and dataset-specific and should not be interpreted as universal compression thresholds or clinical safety boundaries. Overall, the findings support the use of representation-level robustness assessment as a complement to image-level fidelity metrics in AI-oriented digital pathology workflows.

ElectronicsVol. 15(18)
Western University (CA), University of Waterloo (CA)
Openalex Percentile: Top 8%
AI in cancer detection
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.