Advancing emotion recognition: a novel adaptive audio–visual fusion approach
摘要
Human emotions are essential indicators of mental states and are increasingly utilized in applications ranging from healthcare to intelligent systems. This work starts with a concise review of the field, emphasizing key steps for accurate emotion recognition, including feature extraction from both speech and facial modalities and classification using both traditional machine learning and modern deep learning techniques. Major challenges, such as variability across speakers, cultural influences, and the scarcity of annotated datasets, are also discussed. Building on this foundation, the paper introduces an adaptive multimodal fusion model that moves beyond static integration strategies. By leveraging confidence-based weighting and transformer-inspired attention mechanisms, the model dynamically combines audio and visual features, enhancing resilience in scenarios with noisy audio or partially occluded faces. Overall, this study combines a systematic review with a methodological innovation, offering both a synthesis of current knowledge and a practical framework for developing robust and generalizable emotion recognition systems. Experimental results on the RML and SAVEE datasets under multiple evaluation protocols, including stratified hold-out, 5-fold cross-validation, Leave-One-Speaker-Out (LOSO), and cross-dataset evaluation, demonstrate the effectiveness and stability of the proposed framework. The proposed adaptive cross-attention fusion model consistently outperforms unimodal baselines and conventional fusion strategies, while the cross-dataset experiments reveal the remaining challenges of domain generalization in multimodal emotion recognition.