Comparison of Framewise Video Classification in Laryngoscopies
摘要
In this study, the performance of single-task and multi-task models, incorporating both static and temporal classification approaches, for various tasks in medical video laryngoscopy (VL) is assessed through a deep learning (DL) image encoder and LSTM networks. The data foundation is an in-house dataset of 464 individual recordings. In contrast to previousworks,we consider the impact of multi-task learning and temporal dependencies in the data through video snippets. Results show that multi-task models outperform single-task models for tasks with sparse labels, indicating the benefits of shared learning across tasks. Moreover LSTM-based models significantly improve temporal consistency and performance for tasks with inherent temporal dependencies such the process state of VL.