Background <p>Ensuring a standardized level of scoring capability among surgical examiners is crucial for the fairness and validity of competency-based assessments in standardized residency training. However, traditional in-person training models face significant challenges in scalability and accessibility. The aim of this study was to evaluate the “Level-Points” online training model for enhancing scoring accuracy among surgical examiners and identify its key influencing factors.</p> Methods <p>This study included data from the 2024 national surgical examiner training program. The training adopted a model integrating level-based progression and a point-reward system, designed in accordance with the national surgical resident competency assessment standards and covering five dimensions. Both before and after the training, the examiners were required to rate the five dimensions in the clinical scenario videos, with scoring deviation and improvement values as the outcomes.</p> Results <p>A total of 670 surgical examiners were included. The mean dimensional scoring deviation decreased 3.82 points from 19.34 (pre-assessment) to 15.52 (post-assessment) (<i>P</i> &lt; 0.001, adjusted <i>P</i> &lt; 0.001), with statistically significant improvements consistently observed across all five dimensions (All adjusted <i>P</i> &lt; 0.001). Effect sizes ranged from − 0.41 to -1.15 (Cohen’s d), with the overall effect size being large (d = -1.09). For baseline scoring deviations, significant associations were found only with examiners’ specialty background after FDR correction. For training effectiveness, none of the variables showed statistically significant differences after FDR correction. In dimension-specific analyses, four associations remained statistically significant: plastic surgery in the patient encounter (β = -3.345, se = 1.065, <i>P</i> = 0.002, adjusted <i>P</i> = 0.012) and specialty-specific clinical reasoning (β = 3.915, se = 0.896, <i>P</i> &lt; 0.001, adjusted <i>P</i> &lt; 0.001) dimensions, cardiothoracic surgery in the patient encounter dimension (β = -1.996, se = 0.834, <i>P</i> = 0.017, adjusted <i>P</i> = 0.047), and neurosurgery in the communication skills dimension (β = 4.400, se = 0.933, <i>P</i> &lt; 0.001, adjusted <i>P</i> &lt; 0.001). Professional title and geographic region exhibited moderating effects in specific dimensions but did not survive FDR correction. No significant impact on improvement was observed for age or gender (all adjusted <i>P</i> &gt; 0.05).</p> Conclusion <p>The “Level-Points” online training model effectively enhanced scoring accuracy among surgical examiners, demonstrating significant scalability advantages. Specialty background emerged as a key factor influencing both baseline scoring deviation and dimension-specific improvements, highlighting the need for refined, discipline-sensitive calibration strategies within standardized assessment frameworks.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Effectiveness of the “Level-Points” online training model and influencing factors in enhancing scoring accuracy among surgical examiners

  • Lina Zhang,
  • Mengling Yan,
  • Yiwen Qiu,
  • Haolei Cai,
  • Yunxian Yu

摘要

Background

Ensuring a standardized level of scoring capability among surgical examiners is crucial for the fairness and validity of competency-based assessments in standardized residency training. However, traditional in-person training models face significant challenges in scalability and accessibility. The aim of this study was to evaluate the “Level-Points” online training model for enhancing scoring accuracy among surgical examiners and identify its key influencing factors.

Methods

This study included data from the 2024 national surgical examiner training program. The training adopted a model integrating level-based progression and a point-reward system, designed in accordance with the national surgical resident competency assessment standards and covering five dimensions. Both before and after the training, the examiners were required to rate the five dimensions in the clinical scenario videos, with scoring deviation and improvement values as the outcomes.

Results

A total of 670 surgical examiners were included. The mean dimensional scoring deviation decreased 3.82 points from 19.34 (pre-assessment) to 15.52 (post-assessment) (P < 0.001, adjusted P < 0.001), with statistically significant improvements consistently observed across all five dimensions (All adjusted P < 0.001). Effect sizes ranged from − 0.41 to -1.15 (Cohen’s d), with the overall effect size being large (d = -1.09). For baseline scoring deviations, significant associations were found only with examiners’ specialty background after FDR correction. For training effectiveness, none of the variables showed statistically significant differences after FDR correction. In dimension-specific analyses, four associations remained statistically significant: plastic surgery in the patient encounter (β = -3.345, se = 1.065, P = 0.002, adjusted P = 0.012) and specialty-specific clinical reasoning (β = 3.915, se = 0.896, P < 0.001, adjusted P < 0.001) dimensions, cardiothoracic surgery in the patient encounter dimension (β = -1.996, se = 0.834, P = 0.017, adjusted P = 0.047), and neurosurgery in the communication skills dimension (β = 4.400, se = 0.933, P < 0.001, adjusted P < 0.001). Professional title and geographic region exhibited moderating effects in specific dimensions but did not survive FDR correction. No significant impact on improvement was observed for age or gender (all adjusted P > 0.05).

Conclusion

The “Level-Points” online training model effectively enhanced scoring accuracy among surgical examiners, demonstrating significant scalability advantages. Specialty background emerged as a key factor influencing both baseline scoring deviation and dimension-specific improvements, highlighting the need for refined, discipline-sensitive calibration strategies within standardized assessment frameworks.