Surgical Action Triplet Recognition Assisted by Foundation Models-Based Instrument Localization
摘要
In the field of endoscopic video surgical workflow analysis, surgical action triplet recognition is a comprehensive fine-grained surgical activity recognition task. It establishes data associations among instruments, actions, and targets, providing a standardized description of the interactions between instruments and targets. Due to the lack of spatial annotations, existing methods mostly use weakly supervised learning to locate instruments, which prevents the models from fully utilizing the spatial information in surgical videos, thus reducing the accuracy of triplet recognition. To address this challenge, we ingeniously applied the medical Foundation Model MedSAM [14] to generate spatial annotations of the surgical instruments, providing technical support for precise instrument localization. At the same time, we introduced a supervision learning method based on pseudo-labels, incorporating an instrument segmentation task into the triplet recognition task, using the pseudo-labels generated by the medical Foundation Model as the supervisory signal for the segmentation module. This approach enhances the model’s accuracy in locating surgical instruments, thereby improving the overall precision of triplet recognition. We evaluated our model on the CholecT50 dataset, and compared with the baseline model that uses weakly supervised methods for instrument localization, there was a significant improvement in various indicators of action triplet recognition.