Hybrid CNN-ViT Model for Student Engagement Detection in Open Classroom Environments
摘要
E-learning was developed in one of the steps of the education world, thus providing personalized and flexible learning environments. Nevertheless, the puzzle is yet to get the students to participate in an Open Classroom (OCR) setting in a way that is actively involved. They mostly work alone with digital platforms as they always did, and teachers cannot manage the task of organizing collaboration in this virtual room. Therefore, we introduce a hybrid Convolutional Neural Networks (CNN)-Vision Transformers (ViT) model to detect real-time engagement. In particular, CNN layers are exploited for local features, while the ViT network uses global attention mechanisms to grab spatial and contextual clues, e.g., facial expressions, and body postures. This method is superior because it integrates two different types of modules, which can take advantage of the pros in a limited field and a coverage context, and we can thus achieve a deeper observation of student action. The model was taught with the Video-based Student Engagement Measurement Dataset and there was a success of 85% maximum accuracy, as well as precision (78–83%), recall (75–80%), and f-score (76–82%) levels, which is much higher than the regular techniques. In this way, we offer a system that easily expands, can be used for any purpose, and can be easily read and comprehended as an interpretable solution for student engagement monitoring.