To address the issue of low recognition accuracy of students' sitting postures in classroom scenarios caused by multi-scale, multi-posture, and occlusion, this paper proposes a detection model that fuses original skeletal joint features with skeleton image features. The MobileVit model combines the advantages of CNN and Vision Transformer, adjusts its downsampling method, introduces the convolutional pooling parallel downsampling connection structure, and adds the fusion of local and global features. This enables the network to extract more key information about undesirable sitting features while minimizing the loss of input features. The experiments demonstrate that the model attains an average classification accuracy of 93.4% for four types of sitting postures, surpassing that of existing mainstream methods, while maintaining a low implementation cost.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Recognition of Poor Sitting Posture of Students in Classroom Based on Fusion Features

  • Qiang Zhang,
  • Lianqiang Niu,
  • Sen Lin

摘要

To address the issue of low recognition accuracy of students' sitting postures in classroom scenarios caused by multi-scale, multi-posture, and occlusion, this paper proposes a detection model that fuses original skeletal joint features with skeleton image features. The MobileVit model combines the advantages of CNN and Vision Transformer, adjusts its downsampling method, introduces the convolutional pooling parallel downsampling connection structure, and adds the fusion of local and global features. This enables the network to extract more key information about undesirable sitting features while minimizing the loss of input features. The experiments demonstrate that the model attains an average classification accuracy of 93.4% for four types of sitting postures, surpassing that of existing mainstream methods, while maintaining a low implementation cost.