Can students identify AI? – A cross-sectional quantitative study about AI recognition in tablet-based MCQ assessment among fifth-year undergraduate medical students at Saarland University, Germany

Abstract Background Beyond technical feasibility of Artificial Intelligence (AI)-generated multiple-choice questions (MCQs), their educational value in assessment remains unclear. This study aims to evaluate whether students can distinguish between AI-MCQs and National Licensing Exam (NLE) questions in an exam setting and if they align with the curriculum. Methods In this cross-sectional study 119 year five medical students completed a Family Medicine MCQ tablet-based exam. Participants answered 30 AI-generated and 30 NLE MCQs. AI questions were generated from digital learning materials using ChatGPT-4o and Gemini 1.5 Pro. Experts accepted 82% of the questions, with minor edits to retained items and removing those needing major changes. During the exam, students were asked about the question source and whether each item aligned with the course curriculum. Statistical analyses were obtained using Jamovi 2.3.28.0. Results No significant difference in correct attribution of AI or NLE MCQs (t(29.6) = -1.24, p = .225; t(29.6) = 1.18, p = .246) was observed. No significant correlation was found between item difficulty and recognition (τ_b: p = .534). Distractor distributions did not differ across ChatGPT, Google Gemini and NLE (χ 2 (2) = 2.61, p = .271, U p > .05 for all comparisons). Item difficulty did not differ significantly between AI and NLE items; however, exploratory source-specific analysis showed an overall difference among ChatGPT, Google Gemini, and NLE items (χ²(2) = 6.71, p = .035), with a difference between Google Gemini and NLE items ( p = .028). Student-perceived curricular alignment did not differ significantly between AI and NLE or among ChatGPT, Google Gemini, and NLE. Conclusions Students’ recognition did not differ between AI-generated MCQs and NLE MCQs. Easier MCQs are generally perceived as more aligned with the curriculum. AI-MCQs may provide a feasible approach to item drafting within a structured human-review process.

Authors

Institutions

Publication Details

Journal
BMC Medical Education
Published
2026-09-24
DOI
https://doi.org/10.1186/s12909-026-10410-8
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Can students identify AI? – A cross-sectional quantitative study about AI recognition in tablet-based MCQ assessment among fifth-year undergraduate medical students at Saarland University, Germany

Philip Vogt, Johannes Jäger, Fabian Dupont, Sara Volz-Willems et al.
BMC Medical Education
Artificial Intelligence in Healthcare and Education
article

Can students identify AI? – A cross-sectional quantitative study about AI recognition in tablet-based MCQ assessment among fifth-year undergraduate medical students at Saarland University, Germany

Philip Vogt, Johannes Jäger, Fabian Dupont, Sara Volz-Willems, Nadine Wolf, Sandra Jordan
article en

Abstract

Abstract Background Beyond technical feasibility of Artificial Intelligence (AI)-generated multiple-choice questions (MCQs), their educational value in assessment remains unclear. This study aims to evaluate whether students can distinguish between AI-MCQs and National Licensing Exam (NLE) questions in an exam setting and if they align with the curriculum. Methods In this cross-sectional study 119 year five medical students completed a Family Medicine MCQ tablet-based exam. Participants answered 30 AI-generated and 30 NLE MCQs. AI questions were generated from digital learning materials using ChatGPT-4o and Gemini 1.5 Pro. Experts accepted 82% of the questions, with minor edits to retained items and removing those needing major changes. During the exam, students were asked about the question source and whether each item aligned with the course curriculum. Statistical analyses were obtained using Jamovi 2.3.28.0. Results No significant difference in correct attribution of AI or NLE MCQs (t(29.6) = -1.24, p = .225; t(29.6) = 1.18, p = .246) was observed. No significant correlation was found between item difficulty and recognition (τ_b: p = .534). Distractor distributions did not differ across ChatGPT, Google Gemini and NLE (χ 2 (2) = 2.61, p = .271, U p > .05 for all comparisons). Item difficulty did not differ significantly between AI and NLE items; however, exploratory source-specific analysis showed an overall difference among ChatGPT, Google Gemini, and NLE items (χ²(2) = 6.71, p = .035), with a difference between Google Gemini and NLE items ( p = .028). Student-perceived curricular alignment did not differ significantly between AI and NLE or among ChatGPT, Google Gemini, and NLE. Conclusions Students’ recognition did not differ between AI-generated MCQs and NLE MCQs. Easier MCQs are generally perceived as more aligned with the curriculum. AI-MCQs may provide a feasible approach to item drafting within a structured human-review process.

BMC Medical Education
Martin Luther University Halle-Wittenberg (DE), Saarland University (DE)
Quality Education
Openalex Percentile: Top 15%
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.