A review on an AI-driven face robot for human-robot expression interaction
摘要
As artificial intelligence (AI) extends humanoid robots into social domains like education, healthcare, and home, the need for emotional interaction is increasing. Facial expressions, conveying 55% of emotional information, are key to emotional bonding, making realistic-faced humanoid robots—face robots—increasingly essential. This article reviews AI-driven expression interaction technologies in face robots. It first examines the hardware architecture of face robots, then analyzes a “perception-reasoning-generation” framework by comparing traditional and advanced approaches. Traditional methods rely on visual and speech-based emotion recognition, along with discrete or dimensional emotion models, to drive expression generation (e.g., facial movements, eye contact, and lip synchronization) through affective computing. In contrast, advanced approaches leverage multimodal fusion, large language model (LLM) or multimodal large language model (MLLM)-based emotion reasoning, and agent-based planning, memory, and tool use to enhance adaptability, realism, and emotional intelligence. The article also discusses potential application areas, current challenges, and future research directions for face robots. By integrating progress across hardware, algorithms, applications, and open issues, this review lays a comprehensive foundation for the development of empathetic, socially adaptive face robots suitable for complex human environments.