RGB+Skeleton: Cross-Modal Fusion Training for Human Action Recognition
摘要
Single-modality Human Action Recognition (HAR) approaches often fall short in achieving satisfactory performance in recognizing specific actions, attributable to inherent data features. Therefore, we propose a multimodal fusion approach for HAR, rooted in the fusion of skeleton and RGB video. This framework entails training the model separately using both skeleton and RGB video to extract action features intrinsic to each modality. Subsequently, the fusion of skeleton and RGB video is pursued from two key perspectives: classification results and action features. In terms of classification results, we explore and discuss four methods: the confidence-based optimal selection method, the bimodal weighted sum method, and two variations of models reliant on the fusion of classification probability. Regarding action features, we propose a cross-modal fusion training based on action features, operating on both skeleton and RGB video models. This strategy utilizes features from one modality to augment the training of the other, facilitating multi-modality fusion at the feature level. The proposed multimodal fusion strategy is compared to existing methods on two large datasets: NTU RGB+D and NTU RGB+D 120. Experimental results underscore the effectiveness of the propsoed approach.