AI-Driven Speaker Identification: Enabling Human-Centric and Trustworthy Intelligent Systems
摘要
Text-independent speaker identification remains challenging in real-world settings due to noise variability and limited modeling of spatial–temporal speech dependencies. This paper proposes a hybrid CNN–RNN framework that combines convolutional feature extraction with recurrent temporal modeling to enhance robustness. The approach integrates spectrogram-based inputs with dropout regularization, adaptive learning-rate decay, and extensive data augmentation (noise injection, pitch shifting, time stretching) to improve generalization. Evaluated on the VoxCeleb1 dataset, the model achieves 99.9% accuracy in clean conditions and sustains 90.2% accuracy at 5 dB SNR, outperforming standalone CNN, RNN, and LSTM baselines. These results demonstrate strong noise resilience and statistical stability across multiple runs. The proposed system supports reliable deployment in security, forensic analysis, and human–computer interaction, contributing to trustworthy and human-aligned biometric authentication.