Burn Extent and Fitzpatrick Skin Tone Assessment from Clinical Photographs: Systematic and Random Error in Multimodal Large Language Models

Background: Burn extent guides triage, transfer and fluid resuscitation, yet its clinical estimation is imprecise and observer-dependent. Multimodal large language models (MLLMs) process clinical photographs without task-specific training, but their error has rarely been separated into systematic and random components or their performance across skin tones characterized. Methods: Three state-of-the-art MLLMs (Gemini 3.1 Pro, GPT-5.6 Sol, and Fable 5) each assessed 153 burn photographs five times under an identical prompt. The tasks were as follows: burned proportion of the imaged field, against an expert-guided pixel-wise segmentation (tolerance ± 10 percentage points, pp); burned percentage of total body surface area (TBSA), against physician consensus (±2 pp); and binary Fitzpatrick skin tone (FST; light I–III versus dark IV–VI). The first of these was the primary endpoint. Results: The primary endpoint was in the range of 32.5–70.2%, with TBSA at 69.5–77.5%. All models compressed the estimation range (slopes 0.58–0.76, intercepts +10.0 to +22.2 pp); one multiplicative constant per model brought errors differing more than twofold into a 1.4 pp range. Across repeated queries, the median within-image range was 5.0–25.0 pp; averaging the five answers reduced error by only 0.24–2.22 pp. FST accuracy was 83.8–91.9% against a majority-class baseline of 81.0%. Conclusions: Averaging repeated answers removes only the smaller, random component; the larger, systematic one persists and requires calibration against reference data before clinical use can be considered.

Authors

Institutions

Publication Details

Journal
Bioengineering
Published
2026-08-28
DOI
https://doi.org/10.3390/bioengineering13091000
Primary Topic
Burn Injury Management and Outcomes
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Burn Extent and Fitzpatrick Skin Tone Assessment from Clinical Photographs: Systematic and Random Error in Multimodal Large Language Models

Armin Kraus, Gerrit Grieb, Henrik Stelling, Ibrahim Güler
Bioengineering
Burn Injury Management and Outcomes
article

Burn Extent and Fitzpatrick Skin Tone Assessment from Clinical Photographs: Systematic and Random Error in Multimodal Large Language Models

Armin Kraus, Gerrit Grieb, Henrik Stelling, Ibrahim Güler
article en

Abstract

Background: Burn extent guides triage, transfer and fluid resuscitation, yet its clinical estimation is imprecise and observer-dependent. Multimodal large language models (MLLMs) process clinical photographs without task-specific training, but their error has rarely been separated into systematic and random components or their performance across skin tones characterized. Methods: Three state-of-the-art MLLMs (Gemini 3.1 Pro, GPT-5.6 Sol, and Fable 5) each assessed 153 burn photographs five times under an identical prompt. The tasks were as follows: burned proportion of the imaged field, against an expert-guided pixel-wise segmentation (tolerance ± 10 percentage points, pp); burned percentage of total body surface area (TBSA), against physician consensus (±2 pp); and binary Fitzpatrick skin tone (FST; light I–III versus dark IV–VI). The first of these was the primary endpoint. Results: The primary endpoint was in the range of 32.5–70.2%, with TBSA at 69.5–77.5%. All models compressed the estimation range (slopes 0.58–0.76, intercepts +10.0 to +22.2 pp); one multiplicative constant per model brought errors differing more than twofold into a 1.4 pp range. Across repeated queries, the median within-image range was 5.0–25.0 pp; averaging the five answers reduced error by only 0.24–2.22 pp. FST accuracy was 83.8–91.9% against a majority-class baseline of 81.0%. Conclusions: Averaging repeated answers removes only the smaller, random component; the larger, systematic one persists and requires calibration against reference data before clinical use can be considered.

BioengineeringVol. 13(9)
Friedrich-Alexander-Universität Erlangen-Nürnberg (DE), Gemeinschaftskrankenhaus Havelhöhe (DE), Westfälische Hochschule (DE), Berlin Center for Epidemiology and Health Research (DE), Otto-von-Guericke-Universität Magdeburg (DE)
Quality Education
Openalex Percentile: Top 10%
Burn Injury Management and Outcomes
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.