BabaSpeech: A Deep Learning-Based Translation of Sign Language Into Lingala Text and Speech for Deaf-Mute Inclusivity
摘要
Communication is essential for social and educational inclusion, but it remains a challenge for deaf and mute individuals, especially in the Democratic Republic of the Congo (DRC) and neighboring countries where Lingala is widely spoken. Despite advances in artificial intelligence, existing sign language translation systems have focused mainly on languages such as English and French, failing to address the needs of African language speakers. According to the World Federation of the Deaf (WFD), approximately 70 million deaf individuals live worldwide, with 80% residing in developing countries. These individuals use more than 300 distinct sign languages, each with linguistic structures independent of spoken languages. This paper proposes an intelligent sign language translation system for Lingala, leveraging deep learning techniques. The model is built on a fine-tuned EfficientNet-B3 architecture combined with Long Short-Term Memory (LSTM) for temporal gesture modeling, enabling real-time translation of hand movements into text and speech. This approach enhances accessibility for deaf and mute individuals proficient in Lingala but with limited access to formal education. The results show that the system achieves over 91.8% accuracy, demonstrating its effectiveness for integration into educational programs, language training centers, and media platforms. This work contributes to reducing communication barriers and promoting linguistic accessibility for deaf-mute communities in the DRC and neighboring countries.