<p>Surgical video workflow analysis has made intensive development in computer-assisted surgery by combining deep learning models, aiming to enhance surgical scene analysis and decision-making. However, previous research has primarily focused on coarse-grained analysis of surgical videos, e.g., phase recognition, instrument recognition, and triplet recognition that only considers relationships within surgical triplets. In order to provide a more comprehensive fine-grained analysis of surgical videos, this work focuses on accurately identifying triplets &lt;<i>instrument</i>, <i>verb</i>, <i>target</i>&gt; from surgical videos. Specifically, we propose a vision-language deep learning framework that incorporates intra- and inter- triplet modeling, termed I<sup>2</sup>TM, to explore the relationships among triplets and leverage the model understanding of the entire surgical process, thereby enhancing the accuracy and robustness of recognition. Besides, we also develop a new surgical triplet semantic enhancer (TSE) to establish semantic relationships, both intra- and inter-triplets, across visual and textual modalities. Extensive experimental results on surgical video benchmark datasets demonstrate that our approach can capture finer semantics, achieve effective surgical video understanding and analysis, with potential for widespread medical applications.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Surgical video workflow analysis via visual-language learning

  • Pengpeng Li,
  • Xiangbo Shu,
  • Chun-Mei Feng,
  • Yifei Feng,
  • Wangmeng Zuo,
  • Jinhui Tang

摘要

Surgical video workflow analysis has made intensive development in computer-assisted surgery by combining deep learning models, aiming to enhance surgical scene analysis and decision-making. However, previous research has primarily focused on coarse-grained analysis of surgical videos, e.g., phase recognition, instrument recognition, and triplet recognition that only considers relationships within surgical triplets. In order to provide a more comprehensive fine-grained analysis of surgical videos, this work focuses on accurately identifying triplets <instrument, verb, target> from surgical videos. Specifically, we propose a vision-language deep learning framework that incorporates intra- and inter- triplet modeling, termed I2TM, to explore the relationships among triplets and leverage the model understanding of the entire surgical process, thereby enhancing the accuracy and robustness of recognition. Besides, we also develop a new surgical triplet semantic enhancer (TSE) to establish semantic relationships, both intra- and inter-triplets, across visual and textual modalities. Extensive experimental results on surgical video benchmark datasets demonstrate that our approach can capture finer semantics, achieve effective surgical video understanding and analysis, with potential for widespread medical applications.