Resolution-dependent self-supervised transfer in chest radiograph classification
Abstract Background: Self-supervised learning (SSL) has improved visual representation learning, but its value in chest radiography remains uncertain. DINOv3 extends earlier SSL models through Gram-anchored self-distillation and explicit high-resolution adaptation. Whether these changes improve transfer learning for chest radiograph classification has not been established. Methods: We benchmarked DINOv3 against DINOv2 and supervised ImageNet initialization across seven chest radiograph datasets comprising 816,183 radiographs from pediatric and adult cohorts. ViT-B/16 and ConvNeXt-B were evaluated under full fine-tuning at 224 × 224 and 512 × 512 pixels, with targeted 1024 × 1024 experiments on three cohorts. Additional analyses examined parameter-efficient adaptation, synthetic label corruption, external validation, frozen 7B features, and computational efficiency. The primary outcome was the mean area under the receiver operating characteristic curve across labels. Results: In adult cohorts, DINOv3 did not consistently outperform DINOv2 at 224 × 224 pixels, but became the strongest initialization at 512 × 512 pixels, especially with ConvNeXt-B. Gains were greatest for small focal and boundary-dependent abnormalities, whereas large-structure findings changed little. The pediatric cohort showed no significant benefit from DINOv3, higher resolution, or backbone choice. Scaling to 1024 × 1024 rarely improved performance and markedly increased computational cost. ConvNeXt-B remained superior to ViT-B/16 under both full and parameter-efficient adaptation. External validation preserved the 512 × 512 DINOv3 advantage, whereas synthetic label corruption showed that this benefit should not be interpreted simply as superior noise robustness. Frozen DINOv3-7B features underperformed relative to fully adapted 86 to 89M-parameter backbones. Conclusions: For adult chest radiograph classification, DINOv3 provides its most reliable benefit at 512 × 512 pixels, particularly with ConvNeXt-B. Fully adapted mid-sized models at 512 × 512 pixels provided the best performance-cost trade-off in our benchmark.
Authors
- Daniel Truhn (ORCID: https://orcid.org/0000-0002-9605-0728)
- Jakob Nikolas Kather (ORCID: https://orcid.org/0000-0002-3730-5348)
- Mina Shaigan (ORCID: https://orcid.org/0000-0003-1719-9944)
- Soroosh Tayebi Arasteh (ORCID: https://orcid.org/0000-0003-1015-7733)
- Sven Nebelung (ORCID: https://orcid.org/0000-0002-5267-9962)
- Christiane Kuhl
Institutions
- Heidelberg University (DE)
- University Hospital Heidelberg (DE)
- Fresenius (Germany) (DE)
- National Center for Tumor Diseases (DE)
- Else Kröner Fresenius Center for Digital Health (DE)
- RWTH Aachen University (DE)
- Stanford University (US)
Publication Details
- Journal
- Communications Medicine
- Published
- 2026-09-09
- DOI
- https://doi.org/10.1038/s43856-026-01897-9
- Primary Topic
- COVID-19 diagnosis using AI
- Type
- article
- Field-Weighted Citation Impact
- 0.00