Letter to Editor for “Diagnostic Performance of Multimodal Large Language Models in Assigning Bone-RADS Categories on CT”

I have read the article “Diagnostic Performance of Multimodal Large Language Models in Assigning Bone-RADS Categories on CT” written by Kaya et al., which evaluates Large Language Model (LLM) performance in a new era, the Society of Skeletal Radiology Bone-RADS, with great interest. According to the article, it is unclear whether the reference standard depends on histopathological examination. Determining accuracy by treating category 1 as negative and categories 2, 3, and 4 as positive, rather than omitting category 2 and 3 lesions, may have changed the results. In comparison of the three raters, the abdominal radiologist, Open AI ChatGPT, and Gemini, the statistical power was low to detect differences among these three raters. Using trained LLMs for lesions with typical imaging features may provide a fairer comparison. Further research with large sample sizes using trained LLMs may reveal a more realistic comparison, enabling their use in practice, especially for non-musculoskeletal radiologists.

Authors

Institutions

Publication Details

Journal
Uludağ Üniversitesi Tıp Fakültesi Dergisi
Published
2026-10-06
DOI
https://doi.org/10.32708/uutfd.1937411
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Letter to Editor for “Diagnostic Performance of Multimodal Large Language Models in Assigning Bone-RADS Categories on CT”

Dilek Sağlam
Uludağ Üniversitesi Tıp Fakültesi Dergisi
Artificial Intelligence in Healthcare and Education
article

Letter to Editor for “Diagnostic Performance of Multimodal Large Language Models in Assigning Bone-RADS Categories on CT”

Dilek Sağlam
article en

Abstract

I have read the article “Diagnostic Performance of Multimodal Large Language Models in Assigning Bone-RADS Categories on CT” written by Kaya et al., which evaluates Large Language Model (LLM) performance in a new era, the Society of Skeletal Radiology Bone-RADS, with great interest. According to the article, it is unclear whether the reference standard depends on histopathological examination. Determining accuracy by treating category 1 as negative and categories 2, 3, and 4 as positive, rather than omitting category 2 and 3 lesions, may have changed the results. In comparison of the three raters, the abdominal radiologist, Open AI ChatGPT, and Gemini, the statistical power was low to detect differences among these three raters. Using trained LLMs for lesions with typical imaging features may provide a fairer comparison. Further research with large sample sizes using trained LLMs may reveal a more realistic comparison, enabling their use in practice, especially for non-musculoskeletal radiologists.

Uludağ Üniversitesi Tıp Fakültesi DergisiVol. 52
Bursa Uludağ Üni̇versi̇tesi̇ (TR)
Openalex Percentile: Top 18%
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.