Spatio-temporal unsupervised individual clustering for operating room videos
摘要
Human activity recognition (HAR) research has recently focused on multiple individuals within videos. However, conventional models are trained using supervised or semi-supervised learning, which makes their direct application to real-world videos challenging. The purpose of this study is to achieve HAR from real-world videos through completely unsupervised learning. As real-world videos, we target operating room surveillance videos with surgery ongoing. We extract visual features using two autoencoders based on Inception 3D (I3D) and spatial features measured by the L2 norm between the operating table and individuals. These individual features were clustered using a centroid-based model. We evaluated our method on 29 pieces of different operating room videos from 6 different operating rooms in 145 seconds in total, and achieved 0.83 and 0.71 of accuracy in the training and test datasets, respectively, for individual clustering. Our method allows automatic analysis of operating room videos, which contributes to improving the efficiency and effectiveness of postoperative analysis and further medical education.