Multi-Dialect Speech Corpus Creation for Enhancing Tamil Automatic Speech Recognition
摘要
In the current technological revolution, voice assistance systems are widely used. With the use of Automatic Speech Recognition (ASR) technology, a computer can recognize spoken words and translate them into printed text. In general, the spoken form of a language differs from the written form. Conversational systems and their applications have found various applications, such as the operation of various equipment through speech, access to maps for hands-free driving, query response systems for information retrieval, etc. Like human-to-human communication, human–machine communication should ideally be in spoken form to ensure accessibility and usability, accommodating dialect variations so that a broad population can utilize speech-enabled applications seamlessly. The spoken language is now widely used worldwide, but the presence of borrowed or unique words from other languages poses challenges to developing advanced ASR systems. There remains a pressing need for specialized systems capable of accurately recognizing speech in regional languages, particularly to serve underrepresented and underserved populations. Typically, only native speakers can accurately render these spoken forms, which are often not documented in written text. Tamil is a prime example of such a language, with numerous dialects spoken across various regions of Tamil Nadu. Access to digital texts and labeled dialect speech data in Tamil remains scarce, and collecting labeled dialect speech data for the language is a demanding and time-consuming process. This research paper seeks to address this gap by presenting the development and evaluation of a real-time, multi-dialect automatic speech recognition system tailored specifically to the Tamil language, a regional dialect with unique characteristics that pose distinct challenges for conventional ASR technologies. Emerging evidence suggests that the performance of automated speech recognition systems can vary significantly across different demographic groups, with certain subpopulations facing considerable hurdles in effectively utilizing these technologies. To achieve this, we collected dialect-specific Tamil speech data from southern, northern, western, and central regions of Tamil Nadu. Utilizing open-source pre-trained ASR models, we developed a proof-of-concept ASR system. Our data set comprises 24 h and 27 min of Tamil dialect speech data spoken by 240 individuals. We anticipate that our approach will improve opportunities for developing systems capable of accurately recognizing and interacting through spoken language across a diverse range of dialects, speakers, and environmental conditions. This corpus creation supports the training and evaluation of the proposed multi-dialect ASR system.