Recognition of Poor Sitting Posture of Students in Classroom Based on Fusion Features
摘要
To address the issue of low recognition accuracy of students' sitting postures in classroom scenarios caused by multi-scale, multi-posture, and occlusion, this paper proposes a detection model that fuses original skeletal joint features with skeleton image features. The MobileVit model combines the advantages of CNN and Vision Transformer, adjusts its downsampling method, introduces the convolutional pooling parallel downsampling connection structure, and adds the fusion of local and global features. This enables the network to extract more key information about undesirable sitting features while minimizing the loss of input features. The experiments demonstrate that the model attains an average classification accuracy of 93.4% for four types of sitting postures, surpassing that of existing mainstream methods, while maintaining a low implementation cost.