Effectiveness of ChatGPT and DeepSeek in Urology Medical Education: Randomized Controlled Trial
- Journal
- Journal of medical Internet research (Q1)
- Published
- 21 September 2026
- Study design
- Randomized controlled trial
- Evidence level
- Level 1, High (CEBM 1b)
- Authors
- Wentong Yang, Ting Xu, Junjie Wei, Wentao Zheng, Wenbo Yan, Jingkai Wang, et al.
- PMID
- 42766393
- DOI
- 10.2196/89315
Why clinicians should know about it
- Picked for Urology (top studies of the week, 27 September 2026): AI tools for urology education, not patient care
Abstract
BACKGROUND: Since its release in November 2022, generative AI (GenAI) tools, including ChatGPT, have gained widespread attention across various sectors, including medical education. OBJECTIVE: This study seeks to examine the effectiveness and feasibility of GenAI tools (ChatGPT o3‑mini [OpenAI] and DeepSeek R1) in enhancing urology teaching outcomes for medical undergraduates. METHODS: We assessed the accuracy of responses from ChatGPT o3-mini and DeepSeek R1 to authoritative urology multiple-choice questions. Then, a randomized controlled trial was performed to compare the learning outcomes of students using ChatGPT o3-mini and DeepSeek R1 with those using traditional learning methods. Additionally, a questionnaire was designed to survey medical undergraduates' perspectives on the application of AI in urology education. RESULTS: DeepSeek R1 demonstrated higher accuracy than ChatGPT o3-mini in answering urology-related multiple-choice questions. In the test following the self-study period, the DeepSeek R1 group surpassed both the control and ChatGPT o3-mini groups in total scores across various question types. Despite the superior scores in the ChatGPT o3-mini group, statistical significance was not achieved relative to the control group. Survey results revealed that most students had a positive attitude toward AI-assisted learning, believing it could effectively enhance medical education. CONCLUSIONS: DeepSeek-assisted self-study was associated with higher posttest scores than traditional internet-based learning, whereas ChatGPT showed numerically higher but nonsignificant results. These findings offer evidence-based insights into the embedding of GenAI within medical education frameworks, providing guidance for educators in developing teaching strategies and for institutions in formulating relevant policies.
Abstract as published, via PubMed.
For healthcare professionals. The summary is generated by AI from the published abstract, and the evidence level is assigned automatically from the study design on the Oxford CEBM hierarchy. Neither is medical advice. Read the full paper before changing practice.