Emotion Recognition in Multi-speaker Scenarios for Social Robots
摘要
Social robotics is expected to become increasingly integrated into many areas of people’s daily lives. As a result, many research efforts focus on improving Human-Robot Interaction (HRI), including emotion recognition to define robot behavior. This study addresses the challenge of emotion recognition in noisy environments, a critical issue for social robotics, where ambient noise and overlapping voices often hinder the effectiveness of voice-based emotional detection systems. We propose a single channel voice source separation model in real time based on advanced deep learning techniques to enhance voice extraction during conversations. This research also introduces a reference database specifically designed for emotional separation, which is crucial for evaluating the performance of the model relative to previous approaches. We focus on refining both the voice source separation system and the emotion recognition algorithm through continuous optimization and fine-tuning of the models. Our method improves emotion recognition accuracy by 32% compared to a baseline non-separation pipeline, achieving more than 76% accuracy in noisy environments and 63% in particularly challenging scenarios, such as highly reverberant and noisy conditions. Moreover, even when tested in more complex and real scenarios, such as language adaptation and HRI tests, the voice separation performance remains optimal, reinforcing the robustness of our approach. The impact of this work extends beyond social robotics, offering improvements for any application that relies on clear and precise voice recognition. By improving voice separation and reducing background noise, the model is expected to improve HRI in complex situations and contribute to the development of more efficient and reliable emotion recognition systems.