Skip to main content

Randomized double-blind protocol for cross-specialty competency assessment in robotic surgery training using GEARS: a single-center initial validation study

In brief

Blind robotic-skill scoring separates trainees from experts with 76% reliability

In a single-center trial of 49 novices across general surgery, urology and gynecology, a randomized double-blind protocol using the GEARS score showed strong validity (KMO 0.84) and good internal consistency (Cronbach α 0.76), reliably distinguishing novices from expert videos. The tool performed consistently across specialties, models and genders, suggesting it could serve as a standard competency metric, though broader multi-center validation is still needed.

Journal
Journal of robotic surgery (Q1)
Published
20 July 2026
Study design
Randomized controlled trial
Evidence level
Level 1, High (CEBM 1b)
Authors
Huiqin Zhou, Yuxuan Zheng, Yujun Peng, Xuewei Zhu, Lin Zhang, Kun Yang
PMID
42474855
DOI
10.1007/s11701-026-03642-9

Why clinicians should know about it

Abstract

This study aims to evaluate the discriminatory validity of a randomized, double-blind evaluation protocol based on the Global Evaluative Assessment of Robotic Skills (GEARS) scoring system for cross-specialty competency assessment during robotic surgery training. We conducted a prospective, randomized, double-blind evaluation protocol at a single robotic training center. 49 surgical trainees (17 general surgery, 16 urology, 16 gynecology) with no prior robotic experience completed a 3-week training program-7 days of intensive simulator (dV-Trainer) and porcine-model procedures, followed by 2 weeks of clinical observation. All operation videos were anonymized and randomly assigned to 7 blinded expert reviewers per session, selected from a 12-member panel. To calibrate scoring and detect bias, 17 expert-generated videos were randomly interspersed as internal controls. Validity and inter-scenario consistency were assessed using factor analysis, Cronbach's α, and the Bland-Altman method. The GEARS-based assessment protocol under the randomized double-blind design demonstrated good validity (KMO = 0.836, cumulative variance 85%) and reliability (Cronbach's α = 0.765), effectively distinguishing trainees from experts in simulator and live tissue operations (p < 0.05). The Mscore-Sim correlated significantly with the GEARS dimensions (R²=0.563); the Bland-Altman limits of agreement (-2.1-3.8) validated cross-modal scoring consistency. Subgroup analysis revealed the system's stability across specialties (Δ < 0.4), training robot models (Δ < 0.2), and sexes (p > 0.05). Compared to younger trainees, older trainees showed no difference in live tissue performance, suggesting a compensatory effect of experience. The scoring system objectively differentiates operational abilities of trainees across specialties, demonstrating its potential as a standardized assessment tool for heterogeneous robotic surgery training.

Abstract as published, via PubMed.

View on PubMedFull text at the publisherOpen in the app

For healthcare professionals. The summary is generated by AI from the published abstract, and the evidence level is assigned automatically from the study design on the Oxford CEBM hierarchy. Neither is medical advice. Read the full paper before changing practice.