<p>The revolutionary growth in smart devices and increase in social media led to growth of online voice-based applications and online speech communication. The rapid advancement of artificial intelligence and machine learning has significantly impacted speech processing applications, enabling seamless human-computer interaction. This paper presents a Telugu Speech Recognition and Synthesis System employing a hybrid approach integrating Hidden Markov Models (HMM) and Deep Neural Networks (DNN). The system performs Speech-to-Text-to-Speech (STTS) conversion, ensuring accurate speech recognition and synthesis. The proposed model preprocesses Telugu speech data through noise reduction, silence removal, and resampling techniques. Mel-Frequency Cepstral Coefficients (MFCCs) are extracted as feature representations, followed by HMM-based phoneme modelling and DNN-based classification for robust recognition. Experimental results have shown performance in accuracy, computational efficiency, and naturalness of generated speech.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

DNNT (Deep Neural Network for Telugu): a framework for speech recognition of Telugu language with parallel computing approach

  • S. Satheeswara Reddy,
  • Sheik Khadar Ahmad Mnoj,
  • Karu Prasada Rao

摘要

The revolutionary growth in smart devices and increase in social media led to growth of online voice-based applications and online speech communication. The rapid advancement of artificial intelligence and machine learning has significantly impacted speech processing applications, enabling seamless human-computer interaction. This paper presents a Telugu Speech Recognition and Synthesis System employing a hybrid approach integrating Hidden Markov Models (HMM) and Deep Neural Networks (DNN). The system performs Speech-to-Text-to-Speech (STTS) conversion, ensuring accurate speech recognition and synthesis. The proposed model preprocesses Telugu speech data through noise reduction, silence removal, and resampling techniques. Mel-Frequency Cepstral Coefficients (MFCCs) are extracted as feature representations, followed by HMM-based phoneme modelling and DNN-based classification for robust recognition. Experimental results have shown performance in accuracy, computational efficiency, and naturalness of generated speech.