Improving surgical phase recognition using self-supervised deep learning
摘要
In recent decades, there has been growing interest in developing intelligent systems that provide real-time decision support to surgeons in the operating room. Surgical Phase Recognition (SPR) can enhance workflows by monitoring progress and delivering timely feedback. However, SPR development is often constrained by the limited availability of large, labeled surgical videos datasets, due to the costs associated with acquisition and annotation. Self-Supervised Learning (SSL) provides a transformative approach by leveraging unlabeled data to learn robust representations. This study explores the novel application of SSL to SPR in endoscopic pituitary surgery, comparing the performance of two SSL frameworks, SimCLR and BYOL, on a downstream SPR task. An attention-weighted pooling operator is also integrated to enhance spatial feature extraction. Results show two key findings. First, when trained on the full dataset, SimCLR with attention reaches an F1-score of 66% [63%-69%], outperforming the 55% [48%-61%] obtained with fully supervised learning. Second, SSL maintains the quality of the learned image representations even with a 50% reduction in annotated data size, achieving an F1-score of 64% [61%-67%]. Across all evaluations, SimCLR outperformed BYOL, showing greater robustness to intra-class variability. This first application of SSL to endoscopic pituitary surgery shows that Self-Supervised Learning is a robust approach for enhancing Surgical Phase Recognition, particularly when integrated with attention. These findings have important implications for the development of advanced decision support systems in surgery, enabling comparable performance with significantly fewer labeled images.