In this study, the performance of single-task and multi-task models, incorporating both static and temporal classification approaches, for various tasks in medical video laryngoscopy (VL) is assessed through a deep learning (DL) image encoder and LSTM networks. The data foundation is an in-house dataset of 464 individual recordings. In contrast to previousworks,we consider the impact of multi-task learning and temporal dependencies in the data through video snippets. Results show that multi-task models outperform single-task models for tasks with sparse labels, indicating the benefits of shared learning across tasks. Moreover LSTM-based models significantly improve temporal consistency and performance for tasks with inherent temporal dependencies such the process state of VL.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Comparison of Framewise Video Classification in Laryngoscopies

  • Ole Felber,
  • Louis Bellmann,
  • Philipp Breitfeld,
  • Martin Petzoldt,
  • Felix Rindt,
  • René Werner,
  • Maximilian Nielsen

摘要

In this study, the performance of single-task and multi-task models, incorporating both static and temporal classification approaches, for various tasks in medical video laryngoscopy (VL) is assessed through a deep learning (DL) image encoder and LSTM networks. The data foundation is an in-house dataset of 464 individual recordings. In contrast to previousworks,we consider the impact of multi-task learning and temporal dependencies in the data through video snippets. Results show that multi-task models outperform single-task models for tasks with sparse labels, indicating the benefits of shared learning across tasks. Moreover LSTM-based models significantly improve temporal consistency and performance for tasks with inherent temporal dependencies such the process state of VL.