Effectiveness of large language model-based teaching interventions in medical education: a systematic review and stratified meta-analysis of 14 studies
摘要
In recent years, artificial intelligence has been adopted as an innovative educational approach and may substantially reshape health-care delivery systems. This systematic review and stratified meta-analysis compared large language model (LLM)-based teaching interventions with traditional teaching approaches in medical education.
MethodsBased on a four-database search conducted through 31 March 2026, we included comparative studies and feasibility studies evaluating LLM-based teaching tools in medical education learners. This review was not prospectively registered. It evolved from a broader evidence-mapping exercise into a formal systematic review with meta-analysis as the intervention literature and analyzable outcomes became clearer, and no publicly archived protocol was available. Risk of bias was appraised with design-specific RoB 2-, ROBINS-I-, or JBI-informed domains. Domain-specific quantitative synthesis focused on strict knowledge acquisition, delayed retention, and structured clinical performance. An inverse-variance random-effects model was used for pooled SMD analyses. The broader post-course score-based synthesis was retained only as an exploratory summary across related but non-identical outcomes. Studies not directly compatible with a parallel-group SMD framework were retained through classification summary figures and single-study exploratory displays. Certainty of evidence was judged using a GRADE-informed domain-level approach.
ResultsSearches of four databases identified 8103 records (PubMed n = 2203, Web of Science Core Collection n = 2388, Embase n = 1790, and Cochrane Library n = 1722). After removal of 5652 duplicate records and 1251 non-full-text items, 1200 records underwent title/abstract screening, 25 full-text articles were assessed, and 14 studies comprising 1207 analyzable learners were included in the systematic review. Across the domain-specific pooled analyses, structured clinical performance showed the clearest positive signal (2 studies; 38 vs. 39 participants; SMD 3.39, 95% CI 1.43 to 5.35), strict knowledge acquisition showed a smaller but still positive signal (5 studies; SMD 0.80, 95% CI 0.07 to 1.53), and delayed retention showed no clear difference (2 studies; SMD 0.15, 95% CI -0.50 to 0.80). This structured clinical performance estimate was based on a very small evidence base and may be inflated by small-study effects, selective reporting, or measurement artifacts. An exploratory pooled synthesis of post-course score-based outcomes included 7 studies (248 vs. 251 participants) and suggested an advantage for LLM-based teaching (SMD 0.90, 95% CI 0.18 to 1.62, I² = 92.5%, P = 0.015), although heterogeneity across constructs, learner groups, and intervention modes was substantial. Exploratory subgroup analysis suggested a larger effect in practice courses than in theory courses, although this signal was driven by only 2 practice-oriented studies and should be interpreted cautiously. Seven studies did not enter the pooled score-based synthesis directly; 1 was retained in the structured clinical performance supplementary meta-analysis and 6 were preserved through classification summary figures and single-study exploratory displays, thereby allowing all 14 included studies to remain within the overall results framework. GRADE-informed certainty ranged from very low to low across pooled domains.
ConclusionsIn medical education, current evidence provides low-certainty signals that some LLM-based teaching interventions may improve structured clinical performance and certain short-term knowledge outcomes, whereas delayed retention remains uncertain. The broader pooled score-based synthesis should be interpreted only as an exploratory summary across heterogeneous educational constructs, learner groups, and intervention modes, not as a single universal endpoint. Because several pooled domains were based on only 2 to 5 studies and overall certainty remained low to very low, these findings should be regarded as cautious signals rather than definitive estimates.