Image and Speech-Based Emotion Recognition and Gender Identification Using MobileNet and LSTM
摘要
Emotion recognition and gender identification are pivotal in enhancing human–computer interactions, enabling systems to respond more naturally and effectively to user states. This research presents a multi-modal approach integrating both image and speech data to improve the accuracy of emotion recognition and gender classification. Leveraging deep learning architectures, specifically MobileNet for image processing and Long Short-Term Memory (LSTM) networks for speech analysis, the system utilizes comprehensive datasets such as RAVDESS, Mozilla Common Voice, and a Kaggle dataset comprising 30,000 facial images. The proposed system is structured into three primary modules: a Gender Classification Module achieving 91.46% accuracy, an Emotion Recognition Module from speech attaining 86% accuracy, and an Emotion Detection Module from facial expressions to enhance overall recognition precision. By combining these modalities, the system addresses common challenges like noise, variability in speech, and uncontrolled environmental conditions, thereby surpassing the limitations of single-modality systems. The integration of speech and visual data not only improves recognition accuracy but also ensures greater robustness and adaptability across diverse real-world applications, including customer service, healthcare diagnostics, and advanced human–computer interfaces. This study underscores the efficacy of multi-modal deep learning frameworks in creating more empathetic and responsive automated systems, paving the way for future advancements in emotional artificial intelligence.