<p>Dysphonia is a prevalent speech disorder that manifests as irregularities in vocal quality, which presents significant challenges to effective communication. This disorder not only hampers day-to-day interactions but also profoundly affects an individual’s quality of life. Despite its impact, there is a limited focus on addressing dysphonia for non-English languages. This work introduces a methodology for the enhancement of dysphonic speech in Kannada, one of the most widely spoken languages in India. The proposed solution leverages advanced techniques in signal processing and machine learning to restore clarity and intelligibility to dysphonic speech, thereby facilitating better communication for affected individuals. The dataset consists of regularly speaking Kannada sentences recorded from dysphonic subjects in low noise environment. Additive noise present while recording is reduced using spectral subtraction. Noise reduced dysphonic speech is enhanced and reconstructed using deep learning methods like long short-term memory (LSTM) and convolution neural network (CNN). The outcome of the methods is analyzed using quantitative evaluation.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Enhancement and reconstruction of dysphonic Kannada speech using LSTM and convolution network

  • P. Rajeswari,
  • N. Shankaraiah

摘要

Dysphonia is a prevalent speech disorder that manifests as irregularities in vocal quality, which presents significant challenges to effective communication. This disorder not only hampers day-to-day interactions but also profoundly affects an individual’s quality of life. Despite its impact, there is a limited focus on addressing dysphonia for non-English languages. This work introduces a methodology for the enhancement of dysphonic speech in Kannada, one of the most widely spoken languages in India. The proposed solution leverages advanced techniques in signal processing and machine learning to restore clarity and intelligibility to dysphonic speech, thereby facilitating better communication for affected individuals. The dataset consists of regularly speaking Kannada sentences recorded from dysphonic subjects in low noise environment. Additive noise present while recording is reduced using spectral subtraction. Noise reduced dysphonic speech is enhanced and reconstructed using deep learning methods like long short-term memory (LSTM) and convolution neural network (CNN). The outcome of the methods is analyzed using quantitative evaluation.