Multi-level Fusion for Multi-modal Human Action Recognition
摘要
This paper presents a fusion system for human action recognition based on three modalities consisting of depth, skeleton, and inertial data. In order to capture spatiotemporal features, the depth, skeleton, and inertial data are transformed to depth motion map(DMM), skeleton temporal map(STM), and inertial signal image(ISI), respectively. Three-levels fusion methods including data-level, feature-level, and decision-level are taken into consideration. To evaluate the effectiveness and efficiency of the proposed approaches, extensive experiments were conducted on the Changzhou University Multimodal Human Action Dataset (CZU-MHAD), which simultaneously contains depth images, skeleton sequence and inertial signals. Experimental results demonstrate that these fusion approaches achieve higher recognition performance compared to the single modality. And the feature-level fusion method exhibits the best performance compared to the other ways.