Explainable AI for Voice-Based Gender Recognition Using MFCC Features and Data Augmentation: Invited Paper
摘要
Gender recognition by means of human voice is advantageous in various applications. Thus, it presents a rapidly expanding field of research, driven by advances in Machine Learning (ML) that achieved remarkable performance and demonstrated impressive outcomes. Nevertheless, ML methods often lack transparency due to its black-box nature, the case that hampers understanding the internal reasoning actions to converge on precise decisions. This paper addresses the challenge of designing an interpretable ML approach for voice-based gender recognition. In the proposed approach, Mel-Frequency Cepstral Coefficients (MFCC) are extracted as features from the input human voice signals. Additionally, several data augmentation techniques are employed for exploring their impact on the performance of the proposed gender recognition approach. Moreover, in order to gain more comprehensive understanding on how extracted features contribute to model performance, two essential Explainable Artificial Intelligence (XAI) methods, namely, Local Interpretable Model-Agnostic Explanations (LIME) and SHapley Additive exPlanations (SHAP), are incorporated within the proposed approach.