<p>Minimally invasive surgery has numerous advantages compared to open surgery and occupies a significant position in modern surgical practice. Nevertheless, the deficiency of depth information in endoscopic procedures presents a substantial obstacle for surgeons during operation execution. In view of this, many methods have been proposed for depth estimation within the endoscopic field of view. Recently, some foundation model-based methods have been put forward and have attained remarkable success in enhancing depth estimation accuracy. However, the crucial aspects of effectively extracting valuable high-frequency details, as well as the efficient handling of highlighted pixels, have frequently been neglected. In this work, we augment the high-frequency information extraction ability by leveraging the multi-scale attention module. Moreover, we enhance both local and global information extraction by combining the proposed multi-scale large kernel convolution layers and Transformers. Finally, we address the issue of specular highlights by employing the proposed highlight-aware photometric loss. Extensive experiments have validated the effectiveness of the proposed method. On the SCARED dataset, our method outperforms state-of-the-art method by approximately 3.4% in RMSE. On the ColonDepth dataset, our method demonstrates better generalization, achieving a 7.2% improvement in RMSE, and an 11.7% improvement in RMSE log. We believe that better depth estimation results can promote applications such as MIS and naked-eye three-dimensional endoscopy. Our code will be available at <a href="https://github.com/xypeng22/HADepth">https://github.com/xypeng22/HADepth</a>.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

HADepth: Highlight-aware monocular depth estimation for endoscopy

  • Xiaoyuan Peng,
  • Shigang Wang,
  • Jian Wei,
  • Yan Zhao

摘要

Minimally invasive surgery has numerous advantages compared to open surgery and occupies a significant position in modern surgical practice. Nevertheless, the deficiency of depth information in endoscopic procedures presents a substantial obstacle for surgeons during operation execution. In view of this, many methods have been proposed for depth estimation within the endoscopic field of view. Recently, some foundation model-based methods have been put forward and have attained remarkable success in enhancing depth estimation accuracy. However, the crucial aspects of effectively extracting valuable high-frequency details, as well as the efficient handling of highlighted pixels, have frequently been neglected. In this work, we augment the high-frequency information extraction ability by leveraging the multi-scale attention module. Moreover, we enhance both local and global information extraction by combining the proposed multi-scale large kernel convolution layers and Transformers. Finally, we address the issue of specular highlights by employing the proposed highlight-aware photometric loss. Extensive experiments have validated the effectiveness of the proposed method. On the SCARED dataset, our method outperforms state-of-the-art method by approximately 3.4% in RMSE. On the ColonDepth dataset, our method demonstrates better generalization, achieving a 7.2% improvement in RMSE, and an 11.7% improvement in RMSE log. We believe that better depth estimation results can promote applications such as MIS and naked-eye three-dimensional endoscopy. Our code will be available at https://github.com/xypeng22/HADepth.