A LLM-Based Robot Partner with Multi-modal Emotion Recognition
摘要
The integration of Large Language Models (LLMs) with robotic systems has opened new avenues for the development of empathetic and interactive robot partners. This paper introduces a service robot system that incorporates multi-modal emotion recognition and LLM-based emotion dialogue generation. The system captures user emotions through a tri-modal emotion recognition model (TriMER), which processes audio, text, and facial expressions using advanced techniques like BiLSTM, CNN, and Deformable Convolutional Networks (DCN). Experiments conducted using the IEMOCAP dataset show that our TriMER model achieves an accuracy of 74.15% in recognizing emotions. By combining emotion recognition with LLM, the robot can better understand and respond to human emotions, facilitating more natural and empathetic interactions. This development holds promise for applications in elder care, aiming to enhance both physical and mental well-being.