Letter to Editor for “Diagnostic Performance of Multimodal Large Language Models in Assigning Bone-RADS Categories on CT”
I have read the article “Diagnostic Performance of Multimodal Large Language Models in Assigning Bone-RADS Categories on CT” written by Kaya et al., which evaluates Large Language Model (LLM) performance in a new era, the Society of Skeletal Radiology Bone-RADS, with great interest. According to the article, it is unclear whether the reference standard depends on histopathological examination. Determining accuracy by treating category 1 as negative and categories 2, 3, and 4 as positive, rather than omitting category 2 and 3 lesions, may have changed the results. In comparison of the three raters, the abdominal radiologist, Open AI ChatGPT, and Gemini, the statistical power was low to detect differences among these three raters. Using trained LLMs for lesions with typical imaging features may provide a fairer comparison. Further research with large sample sizes using trained LLMs may reveal a more realistic comparison, enabling their use in practice, especially for non-musculoskeletal radiologists.
Authors
- Dilek Sağlam (ORCID: https://orcid.org/0000-0002-5778-6847)
Institutions
- Bursa Uludağ Üni̇versi̇tesi̇ (TR)
Publication Details
- Journal
- Uludağ Üniversitesi Tıp Fakültesi Dergisi
- Published
- 2026-10-06
- DOI
- https://doi.org/10.32708/uutfd.1937411
- Primary Topic
- Artificial Intelligence in Healthcare and Education
- Type
- article
- Field-Weighted Citation Impact
- 0.00