Automated grading of short-answer image-based assessments using a hybrid natural language processing–AI framework: validation in radiology education

INTRODUCTION: Efficient and reliable assessment is essential in medical education. Manual grading of short-answer radiology quizzes is resource-intensive and variable. Automated marking systems can alleviate these challenges. This study introduces a novel automated grading system leveraging techniques in natural language processing (NLP) and generative artificial intelligence. Our objective was to assess the system's agreement against human grading and to analyse how response length influences grading concordance. METHODS: We obtained 840 responses from 28 radiology residents to 30 X-ray quiz questions. The system employed string-matching algorithms and rule-based logic for anatomical and laterality mismatches. A ChatGPT-created synonym dictionary was used to recognise acceptable alternative answers. Cohen's κ was used to assess inter-rater agreement across answer lengths, and logistic regression was used to examine the response length's relationship with grading disagreement. RESULTS: To establish a reference standard, two independent human raters graded the responses, achieving near-perfect inter-rater reliability (κ = 0.985). Evaluated against human consensus, the system achieved 97.5% raw agreement (816/837). Perfect agreement (κ = 1.00) was observed for one-word responses (n = 429), with near-perfect agreement for two-word responses (κ = 0.927, n = 75). Agreement decreased as response length increased, with each additional word increasing the odds of disagreement by 51.9% (odds ratio 1.519; 95% confidence interval 1.227-1.879). CONCLUSION: The automated grading system, enhanced by NLP techniques and ChatGPT-assisted synonym expansion, demonstrated strong concordance with human grading, particularly for one- and two-word answers. This hybrid approach offers a scalable solution for short-answer assessment, with strong potential in radiology and other medical education contexts.

Authors

Institutions

Publication Details

Journal
Singapore Medical Journal
Published
2026-09-25
DOI
https://doi.org/10.4103/singaporemedj.smj-2025-302
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Automated grading of short-answer image-based assessments using a hybrid natural language processing–AI framework: validation in radiology education

Win Thida, Phua Hwee Tang, Xu Hao Isaac Tan, Zhen Li Samantha Lee
Singapore Medical Journal
Artificial Intelligence in Healthcare and Education
article

Automated grading of short-answer image-based assessments using a hybrid natural language processing–AI framework: validation in radiology education

Win Thida, Phua Hwee Tang, Xu Hao Isaac Tan, Zhen Li Samantha Lee
article en

Abstract

INTRODUCTION: Efficient and reliable assessment is essential in medical education. Manual grading of short-answer radiology quizzes is resource-intensive and variable. Automated marking systems can alleviate these challenges. This study introduces a novel automated grading system leveraging techniques in natural language processing (NLP) and generative artificial intelligence. Our objective was to assess the system's agreement against human grading and to analyse how response length influences grading concordance. METHODS: We obtained 840 responses from 28 radiology residents to 30 X-ray quiz questions. The system employed string-matching algorithms and rule-based logic for anatomical and laterality mismatches. A ChatGPT-created synonym dictionary was used to recognise acceptable alternative answers. Cohen's κ was used to assess inter-rater agreement across answer lengths, and logistic regression was used to examine the response length's relationship with grading disagreement. RESULTS: To establish a reference standard, two independent human raters graded the responses, achieving near-perfect inter-rater reliability (κ = 0.985). Evaluated against human consensus, the system achieved 97.5% raw agreement (816/837). Perfect agreement (κ = 1.00) was observed for one-word responses (n = 429), with near-perfect agreement for two-word responses (κ = 0.927, n = 75). Agreement decreased as response length increased, with each additional word increasing the odds of disagreement by 51.9% (odds ratio 1.519; 95% confidence interval 1.227-1.879). CONCLUSION: The automated grading system, enhanced by NLP techniques and ChatGPT-assisted synonym expansion, demonstrated strong concordance with human grading, particularly for one- and two-word answers. This hybrid approach offers a scalable solution for short-answer assessment, with strong potential in radiology and other medical education contexts.

Singapore Medical Journal
SingHealth Duke-NUS Academic Medical Centre (SG), KK Women's and Children's Hospital (SG)
Quality Education
Openalex Percentile: Top 15%
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.