Cataract surgery is a complex procedure requiring precise execution of multiple steps. To improve the accuracy and efficiency of cataract surgeries, we present a semantic segmentation model for cataract surgery scenes. Our model leverages Unsupervised Domain Adaptation (UDA) techniques to enhance segmentation performance in clinical surgical environments, addressing challenges such as domain shift, occlusions between surgical tools and tissues, and long tail problem. The model utilizes a Teacher-Student model, where a student model is trained in the target domain with pseudo-labels generated by an Exponential Moving Average (EMA) teacher model, ensuring robust learning across domains. Additionally, we utilize a Masked Image Consistency (MIC) module to improve the model’s understanding of occluded regions by enforcing consistency between masked and unmasked predictions. To mitigate class imbalance between anatomical structures and surgical tools, we employ a maximum squares loss, enabling the model to achieve balanced learning. Our results demonstrate that the proposed model improves segmentation accuracy and robustness in cataract surgery scenarios.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Domain Adaptation for Semantic Segmentation of Cataract Surgical Images Based on Masked Image Consistency

  • Yuzhu Zhang,
  • Yijie Pan,
  • Mingyang Ou,
  • Guanghui Gong,
  • Qinhu Zhang,
  • Haojin Li,
  • Heng Li

摘要

Cataract surgery is a complex procedure requiring precise execution of multiple steps. To improve the accuracy and efficiency of cataract surgeries, we present a semantic segmentation model for cataract surgery scenes. Our model leverages Unsupervised Domain Adaptation (UDA) techniques to enhance segmentation performance in clinical surgical environments, addressing challenges such as domain shift, occlusions between surgical tools and tissues, and long tail problem. The model utilizes a Teacher-Student model, where a student model is trained in the target domain with pseudo-labels generated by an Exponential Moving Average (EMA) teacher model, ensuring robust learning across domains. Additionally, we utilize a Masked Image Consistency (MIC) module to improve the model’s understanding of occluded regions by enforcing consistency between masked and unmasked predictions. To mitigate class imbalance between anatomical structures and surgical tools, we employ a maximum squares loss, enabling the model to achieve balanced learning. Our results demonstrate that the proposed model improves segmentation accuracy and robustness in cataract surgery scenarios.