Inter- and Intra-Rater Reliability of a Standardized Infrared Thermographic Cellulite Severity Classification System: A Multi-Rater Study

Background/Objectives: Clinical cellulite grading relies on inspection and palpation and is susceptible to examiner variability. A single experienced assessor previously evaluated a five-grade infrared thermographic classification, but its reproducibility across raters was unknown. To assess inter-rater and intra-rater reliability of a standardized thermographic grading workflow applied by trained novice raters. Methods: This secondary reliability study included 78 anterior FLIR P620 scans. Five raters analyzed each scan twice, one week apart, using standardized FLIR Tools settings and a ΔT-based 0–IV scale. The prespecified principal agreement analysis used Gwet’s AC2 with quadratic weights across the full 0–IV scale; a sensitivity analysis recalculated AC2 using only the observed categories. Confidence intervals were estimated by scan-level bootstrap. Results: The scans yielded 780 ratings, all within grades 0–II. Full five-rater exact agreement was 26.9% and 33.3% in rounds 1 and 2, respectively, while within-one-grade agreement was 89.7% and 92.3%. Observed-category AC2, which restricts quadratic weighting to the grades represented in the data (0–II), was 0.786 (95% CI 0.730–0.840) and 0.845 (0.794–0.885). The corresponding prespecified full-scale AC2 values calculated across the theoretical 0–IV weighting range were 0.949 (0.936–0.962) and 0.963 (0.951–0.972). Intra-rater exact agreement ranged from 66.7% to 97.4%, with 100% agreement within one grade for all raters. Conclusions: The workflow demonstrated reproducible ordinal grading within the lower severity categories represented in the dataset, although exact five-rater agreement remained limited. The findings do not establish reliability for grades III–IV, diagnostic accuracy, or clinical utility.

Authors

Institutions

Publication Details

Journal
Journal of Clinical Medicine
Published
2026-09-20
DOI
https://doi.org/10.3390/jcm15187305
Primary Topic
Infrared Thermography in Medicine
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Inter- and Intra-Rater Reliability of a Standardized Infrared Thermographic Cellulite Severity Classification System: A Multi-Rater Study

Magdalena Głowacka, Szczepańska Patrycja, Barbara Algiert‐Zielińska, Andrzej M. Śliwczyński et al.
Journal of Clinical Medicine
Infrared Thermography in Medicine
article

Inter- and Intra-Rater Reliability of a Standardized Infrared Thermographic Cellulite Severity Classification System: A Multi-Rater Study

Magdalena Głowacka, Szczepańska Patrycja, Barbara Algiert‐Zielińska, Andrzej M. Śliwczyński, Anna Erkiert‐Polguj, Wojciech Glinkowski, Aleksandra Rybak, Agata Kmita, Maria Koserczyk
article en

Abstract

Background/Objectives: Clinical cellulite grading relies on inspection and palpation and is susceptible to examiner variability. A single experienced assessor previously evaluated a five-grade infrared thermographic classification, but its reproducibility across raters was unknown. To assess inter-rater and intra-rater reliability of a standardized thermographic grading workflow applied by trained novice raters. Methods: This secondary reliability study included 78 anterior FLIR P620 scans. Five raters analyzed each scan twice, one week apart, using standardized FLIR Tools settings and a ΔT-based 0–IV scale. The prespecified principal agreement analysis used Gwet’s AC2 with quadratic weights across the full 0–IV scale; a sensitivity analysis recalculated AC2 using only the observed categories. Confidence intervals were estimated by scan-level bootstrap. Results: The scans yielded 780 ratings, all within grades 0–II. Full five-rater exact agreement was 26.9% and 33.3% in rounds 1 and 2, respectively, while within-one-grade agreement was 89.7% and 92.3%. Observed-category AC2, which restricts quadratic weighting to the grades represented in the data (0–II), was 0.786 (95% CI 0.730–0.840) and 0.845 (0.794–0.885). The corresponding prespecified full-scale AC2 values calculated across the theoretical 0–IV weighting range were 0.949 (0.936–0.962) and 0.963 (0.951–0.972). Intra-rater exact agreement ranged from 66.7% to 97.4%, with 100% agreement within one grade for all raters. Conclusions: The workflow demonstrated reproducible ordinal grading within the lower severity categories represented in the dataset, although exact five-rater agreement remained limited. The findings do not establish reliability for grades III–IV, diagnostic accuracy, or clinical utility.

Journal of Clinical MedicineVol. 15(18)
Medical University of Warsaw (PL), University of Economics and Innovation (PL), Medical University of Lodz (PL)
Quality Education
Openalex Percentile: Top 11%
Infrared Thermography in Medicine
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.