Visual steatosis assessment: method‐dependent variation and systematically higher estimates compared with AI measurement

Aims Despite its clinical importance, histological steatosis assessment remains poorly standardized and highly variable among pathologists. We examined determinants of this variability by comparing pathologist estimates with quantitative artificial intelligence (AI) based measurements. Methods and Results Ten experienced liver pathologists from four institutions evaluated a limited set of 20 permanent H&E‐stained whole‐slide images (WSIs) from metabolic dysfunction‐associated steatotic liver disease (MASLD) cases (10 core biopsies, 7 wedge biopsies and 3 resections) and 6 additional 500 × 500‐μm image fields from separate permanent H&E‐stained cases selected to represent varying droplet compositions. The WSIs spanned the full spectrum of steatosis severity. Assessments used four phases: default method, steatosis proportionate area (SPA), percentage of hepatocytes containing fat (%HCF) and Banff large‐droplet criteria. Image fields were also assessed using no size cut‐off or cut‐offs based on hepatocyte, nuclear or 2–3× nuclear size. AI models quantified SPA, %HCF and droplet size for comparison with pathologists' assessments. Pathologists differed substantially in terminology, droplet‐size definitions and quantification methods. Method choice (SPA vs. %HCF) was a major determinant of estimate dispersion and discrepancy from AI measurements. Visual estimates exceeded AI quantification by 1.4–3.6×. Pathologists' SPA estimates were increasingly higher than the corresponding AI measurements with increasing burden of droplets with area below 50 μm 2 , whereas, in exploratory analysis, %HCF estimates remained relatively stable relative to the corresponding AI measurements. Mixed‐effects calibration equations were derived to relate visual and digital assessment scales. Conclusion Steatosis assessment lacks standardized terminology and measurement practices. SPA and %HCF are conceptually distinct, and their interchangeable use amplifies discrepancies, particularly in small‐droplet‐rich cases. These findings support standardized, AI‐compatible quantification; the exploratory conversion equations require independent validation before clinical application.

Authors

Institutions

Publication Details

Journal
Histopathology
Published
2026-09-04
DOI
https://doi.org/10.1111/his.70272
Primary Topic
Liver Disease Diagnosis and Treatment
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Visual steatosis assessment: method‐dependent variation and systematically higher estimates compared with AI measurement

Rofyda Elhalaby, Abdelrahman Shaaban, Daniela Allende, Adilson Costa et al.
Histopathology
Liver Disease Diagnosis and Treatment
article

Visual steatosis assessment: method‐dependent variation and systematically higher estimates compared with AI measurement

Rofyda Elhalaby, Abdelrahman Shaaban, Daniela Allende, Adilson Costa, Zong-Ming Eric Chen, Maxwell L. Smith, Marcela Salomao, Hee Eun Lee, Rondell P. Graham, Rish K. Pai, Christopher Hartley, Byoung Uk Park, Priyadharshini Sivasubramaniam, Kristina A. Matkowskyj, Roger K. Moreira, Bashar Hasan, Murli Krishna, Ameya Patil, Lindsey Smith, Lucas Stetzik, Saadiya Nazli
article en

Abstract

Aims Despite its clinical importance, histological steatosis assessment remains poorly standardized and highly variable among pathologists. We examined determinants of this variability by comparing pathologist estimates with quantitative artificial intelligence (AI) based measurements. Methods and Results Ten experienced liver pathologists from four institutions evaluated a limited set of 20 permanent H&E‐stained whole‐slide images (WSIs) from metabolic dysfunction‐associated steatotic liver disease (MASLD) cases (10 core biopsies, 7 wedge biopsies and 3 resections) and 6 additional 500 × 500‐μm image fields from separate permanent H&E‐stained cases selected to represent varying droplet compositions. The WSIs spanned the full spectrum of steatosis severity. Assessments used four phases: default method, steatosis proportionate area (SPA), percentage of hepatocytes containing fat (%HCF) and Banff large‐droplet criteria. Image fields were also assessed using no size cut‐off or cut‐offs based on hepatocyte, nuclear or 2–3× nuclear size. AI models quantified SPA, %HCF and droplet size for comparison with pathologists' assessments. Pathologists differed substantially in terminology, droplet‐size definitions and quantification methods. Method choice (SPA vs. %HCF) was a major determinant of estimate dispersion and discrepancy from AI measurements. Visual estimates exceeded AI quantification by 1.4–3.6×. Pathologists' SPA estimates were increasingly higher than the corresponding AI measurements with increasing burden of droplets with area below 50 μm 2 , whereas, in exploratory analysis, %HCF estimates remained relatively stable relative to the corresponding AI measurements. Mixed‐effects calibration equations were derived to relate visual and digital assessment scales. Conclusion Steatosis assessment lacks standardized terminology and measurement practices. SPA and %HCF are conceptually distinct, and their interchangeable use amplifies discrepancies, particularly in small‐droplet‐rich cases. These findings support standardized, AI‐compatible quantification; the exploratory conversion equations require independent validation before clinical application.

Histopathology
University of Minnesota (US), Cleveland Clinic (US), Mayo Clinic (US), University of Toronto (CA), Medical College of Wisconsin (US), Jacksonville College (US), University of Alabama at Birmingham (US), University of Minnesota Rochester (US), Mayo Clinic in Arizona (US), Mayo Clinic in Florida (US), University of Virginia (US)
Openalex Percentile: Top 10%
Liver Disease Diagnosis and Treatment
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.