Deep Learning-Based Speaker Identification for Individuals with Voice Disorders
摘要
Speaker identification is crucial for recognizing individuals based on their voice, a fundamental mode of human communication that enables the expression of thoughts, emotions, and ideas. Voice characteristics can deviate from normal acoustic patterns due to various physiological, pathological, or psychological conditions, including Parkinson’s disease (PD). With the growing use of speech in human–machine interactions and the need for effective audio management in multimedia applications, understanding and identifying these variations is essential. This study investigates speaker identification for individuals with PD through the analysis of audio data. Specifically, it involves extracting 34 distinct features from 400 voice recordings, encompassing spectral, temporal, and chromatic characteristics, which are the predominant acoustic attributes of the audio samples. These features were chosen due to their relevance in reflecting the vocal impairments associated with PD. By employing deep learning models, specifically convolutional neural networks (CNNs) and artificial neural networks (ANNs), the research demonstrates the effectiveness of these methods in identifying speakers with PD. The results show high accuracy rates of 97.50% for CNNs and 96.25% for ANNs, indicating the reliability of these approaches in distinguishing unique audio characteristics associated with PD-affected voices. These findings suggest significant potential for developing automated diagnostic tools and enhancing human–machine interaction systems tailored to individuals with PD. Future work could focus on validating these models with larger, more diverse datasets, comparing them with existing methods, and exploring additional features to further improve identification accuracy.