Text-to-Speech Conversion for Gujarati Language Using Deep Learning Technique
摘要
Speech serves as the primary and inherent mode of communication among individuals within the human species. Over the past three decades, there has been a concerted effort by humans to develop computers capable of comprehending and engaging in conversations that mimic human speech patterns. In this pursuit, Text-to-Speech (TTS) synthesis has emerged as a pivotal method, involving the transformation of natural language text into audible speech. However, within the realm of scholarly research and academic production, there is a notable gap, particularly in the field of Gujarati studies. This research seeks to fill this void by providing an efficient and accurate TTS conversion, specifically tailored for the Gujarati language. The methodology employed in this study revolves around deep learning-based procedures for voice synthesis. By leveraging a substantial number of text-to-speech pairings, the approach aims to create effective feature demonstrations, essentially bridging the gap between written text and spoken language and thereby better describing event features. Mel Frequency Cepstral Coefficient (MFCC) analysis is used in this procedure. By employing MFCC, the acoustic features inherent in a spoken stream are extracted and utilized in the synthesis process. The paper employs a convolutional neural network (CNN) integrated with a blend of features such as MFCC, along with a softMax classifier, resulting in superior accuracy when contrasted with current systems.