Multimodal Deep Fusion with Hybrid Temporal Attention Networks for Enhanced Academic Performance Prediction and Behavior Analysis
摘要
This study presents two deep learning architectures for multimodal educational analytics: the Multimodal Deep Fusion Network (MDFN) and the Hybrid Temporal Attention Network (HybTAN). Heterogeneous sources were included in MDFN-from wearable sensor streams, academic records, and socioeconomic indicators-together with specific modalities and hierarchical attention-based fusion. HybTAN models temporal dependencies through the integration of CNN-LSTM and performs calibrated uncertainty estimation using Bayesian inference. Tested on the StudentLife dataset boosted by valid socioeconomic data, MDFN and HybTAN ended up having an impressive amount of performance gains against state-of-the-art baselines in CGPA prediction (MSE 0.174), behavioral classification (accuracy 90.5%), stress level regression, and dropout risk detection, illustrated with attention weight visualizations into interpretable insights modality and temporal relevance possible targeted interventions in educational settings.