Applications, advantages, and limitations of artificial intelligence in medical education assessment: a systematic review
Artificial intelligence (AI) is rapidly being integrated into medical education, where assessments are high-stakes and intricately linked to clinical competence, professional judgment, and patient safety. Although several recent reviews have broadly explored the use of AI in medical education, a focused, methodologically transparent synthesis of AI applications in assessment and evaluation remains limited. This systematic review aimed to (i) identify and classify AI tools and methods used in medical education assessment; (ii) examine their advantages, limitations, and ethical implications; (iii) map AI use to established assessment frameworks (Assessment of Learning [AoL], Assessment for Learning [AfL], Assessment as Learning [AaL], validity, and reliability frameworks); and (iv) clarify the evolving role of medical educators. A systematic review was conducted following the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) 2020 guidelines [1]. Six electronic databases (PubMed/MEDLINE, Scopus, Web of Science, Embase, ERIC, and Cochrane Library) were searched for English-language records published between January 2019 and February 2026. The search combined terms related to artificial intelligence, generative AI, machine learning, assessment, evaluation, and medical education. Two reviewers independently screened the titles, abstracts, and full texts against pre-specified PICOS criteria. The methodological quality was appraised using the Medical Education Research Study Quality Instrument (MERSQI). The findings were thematically synthesized. After deduplication, 1,247 records were screened; 187 full-text articles were assessed for eligibility, and 62 studies met the inclusion criteria. Six themes emerged: (i) AI-supported assessment tools and platforms; (ii) AI use across AoL, AfL, and AaL; (iii) validity, reliability, and feasibility considerations; (iv) medical-education–specific applications including written examinations, objective structured clinical examinations (OSCEs – standardized, station-based assessments of clinical competence), simulation-based assessments, and workplace-based assessments (WBAs); (v) limitations including algorithmic bias, transparency, equity, and data privacy; and (vi) the redefined role of medical educators. The strength of the available evidence, judged descriptively on the basis of MERSQI scores, was variable: most studies were small-scale feasibility or proof-of-concept reports, and few were high-quality validation studies. AI has the potential to improve the efficiency, individualization, and timeliness of feedback in medical education assessment, particularly in large-scale written examinations, OSCE grading, and longitudinal competency tracking. However, evidence of improved validity or reliability is limited to narrowly bounded tasks, current evidence is heterogeneous and dominated by feasibility studies, and concerns regarding bias, transparency, equity, and the irreplaceable role of professional human judgment persist. AI should be positioned as a complement, not a substitute, for expert human evaluation. Medical educators, institutions, and regulators must develop validation, governance, and AI literacy frameworks before AI tools can be safely deployed in high-stakes assessments.
Authors
- Engın Karadağ (ORCID: https://orcid.org/0000-0002-9723-3833)
- Murat Polat (ORCID: https://orcid.org/0000-0001-5851-2322)
Institutions
- Khazar University (AZ)
- Anadolu University (TR)
- Akdeniz University (TR)
- Özyeğin University (TR)
Publication Details
- Journal
- BMC Medical Education
- Published
- 2026-10-05
- DOI
- https://doi.org/10.1186/s12909-026-10545-8
- Primary Topic
- Artificial Intelligence in Healthcare and Education
- Type
- article
- Field-Weighted Citation Impact
- 0.00