Towards development of the first continuous speech recognition system in Indian language Nagpuri
摘要
Nagpuri is the most widely used language spoken by the rural people of the Chotanagpur Plateau region of India. Around 6 million people speak Nagpuri as their first language, and around 8 million as their second language. As this is primarily a spoken language, developing an automatic speech recognition (ASR) system in Nagpuri is extremely important. When we searched the literature, we found no open ASR system or relevant resources in Nagpuri. We developed a Nagpuri speech corpus of around 20 hours, employing 53 native speakers. We have experimented with various deep learning architectures, including Convolu- tional Neural Networks (CNN), Recurrent Neural Networks (RNN), Transformer, and Conformer-based networks. As the size of the training data is not suffi- cient, we investigated whether data augmentation can improve the performance. We applied time-stretching and pitch shift operations to enhance the training data. When we utilized the augmented data to train the system, we found a considerable performance improvement.