Effectiveness of ChatGPT and DeepSeek in Urology Medical Education: Randomized Controlled Trial

Abstract Background Since its release in November 2022, generative AI (GenAI) tools, including ChatGPT, have gained widespread attention across various sectors, including medical education. Objective This study seeks to examine the effectiveness and feasibility of GenAI tools (ChatGPT o3‑mini [OpenAI] and DeepSeek R1) in enhancing urology teaching outcomes for medical undergraduates. Methods We assessed the accuracy of responses from ChatGPT o3-mini and DeepSeek R1 to authoritative urology multiple-choice questions. Then, a randomized controlled trial was performed to compare the learning outcomes of students using ChatGPT o3-mini and DeepSeek R1 with those using traditional learning methods. Additionally, a questionnaire was designed to survey medical undergraduates’ perspectives on the application of AI in urology education. Results DeepSeek R1 demonstrated higher accuracy than ChatGPT o3-mini in answering urology-related multiple-choice questions. In the test following the self-study period, the DeepSeek R1 group surpassed both the control and ChatGPT o3-mini groups in total scores across various question types. Despite the superior scores in the ChatGPT o3-mini group, statistical significance was not achieved relative to the control group. Survey results revealed that most students had a positive attitude toward AI-assisted learning, believing it could effectively enhance medical education. Conclusions DeepSeek-assisted self-study was associated with higher posttest scores than traditional internet-based learning, whereas ChatGPT showed numerically higher but nonsignificant results. These findings offer evidence-based insights into the embedding of GenAI within medical education frameworks, providing guidance for educators in developing teaching strategies and for institutions in formulating relevant policies.

Authors

Publication Details

Journal
Journal of Medical Internet Research
Published
2026-09-21
DOI
https://doi.org/10.2196/89315
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Effectiveness of ChatGPT and DeepSeek in Urology Medical Education: Randomized Controlled Trial

Xu Ting, Wenbo Yan, Junjie Wei, Guangdi Chu et al.
Journal of Medical Internet Research
Artificial Intelligence in Healthcare and Education
article

Effectiveness of ChatGPT and DeepSeek in Urology Medical Education: Randomized Controlled Trial

Xu Ting, Wenbo Yan, Junjie Wei, Guangdi Chu, Haitao Niu, Jingkai Wang, Wentao Zheng, Wentong Yang
article en

Abstract

Abstract Background Since its release in November 2022, generative AI (GenAI) tools, including ChatGPT, have gained widespread attention across various sectors, including medical education. Objective This study seeks to examine the effectiveness and feasibility of GenAI tools (ChatGPT o3‑mini [OpenAI] and DeepSeek R1) in enhancing urology teaching outcomes for medical undergraduates. Methods We assessed the accuracy of responses from ChatGPT o3-mini and DeepSeek R1 to authoritative urology multiple-choice questions. Then, a randomized controlled trial was performed to compare the learning outcomes of students using ChatGPT o3-mini and DeepSeek R1 with those using traditional learning methods. Additionally, a questionnaire was designed to survey medical undergraduates’ perspectives on the application of AI in urology education. Results DeepSeek R1 demonstrated higher accuracy than ChatGPT o3-mini in answering urology-related multiple-choice questions. In the test following the self-study period, the DeepSeek R1 group surpassed both the control and ChatGPT o3-mini groups in total scores across various question types. Despite the superior scores in the ChatGPT o3-mini group, statistical significance was not achieved relative to the control group. Survey results revealed that most students had a positive attitude toward AI-assisted learning, believing it could effectively enhance medical education. Conclusions DeepSeek-assisted self-study was associated with higher posttest scores than traditional internet-based learning, whereas ChatGPT showed numerically higher but nonsignificant results. These findings offer evidence-based insights into the embedding of GenAI within medical education frameworks, providing guidance for educators in developing teaching strategies and for institutions in formulating relevant policies.

Journal of Medical Internet ResearchVol. 28
Openalex Percentile: Top 15%
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.