Predictive Analytics of Detection of Speech Disfluency Using GAN
摘要
Stuttering, along with other speech disfluencies, affects millions of people worldwide, with a large proportion experiencing mild stutters during heated conversations. This speech condition interrupts communication fluency by causing involuntary repetitions, elongations, or quiet pauses. It often appears in childhood, affecting roughly 2.5% of children under the age of five. Extensive studies in Western nations, notably in British and American English environments, have investigated the frequency of such problems among youngsters. This research also concentrated on identifying and comprehending voice disfluencies within the context of these linguistic changes. India has one of the world's largest Telugu-speaking populations, with almost 75 million people living in Andhra Pradesh and Telangana. Given this large population, investigating differences within the Telugu language is critical. To address volume restrictions in datasets, data augmentation techniques are used, which try to artificially enlarge the training set by modifying existing data. WaveGAN, or waveform generative adversarial network, is a well-known augmentation approach. Using WaveGAN, more training data can be generated, improving the accuracy of audio classification tasks. Various classification techniques, such as logistic regression (LR), random forest (RF), decision tree (DT), K-nearest neighbors (KNN), AdaBoost, and XGBoost, were used to classify audio samples as fluent or diffluent. Furthermore, various assessment measures were used to assess the performance of these machine learning algorithms.