Deep Learning-Based Assessment of Impersonated Speech Using Score-Level Fusion of Augmented D-Vector Embeddings
摘要
This study aims to evaluate the quality of mimicked speech by integrating spectral and prosodic speaker embeddings to identify the best mimicked artist, assessed via Mean Opinion Score testing. Augmented d-vector embeddings are constructed separately from prosodic features (loudness, pitch, speaking rate, shimmer, tempogram ratio) and spectral features (Mel Frequency Cepstral Coefficients, chroma, tonnetz, spectral roll off, bandwidth, flux, centroid). Each embedding type is processed and classified independently using a Deep Neural Network classifier. The classifiers produce prediction probability scores for all candidate artists, which are then fused at the score level using empirically chosen weighting constants to generate a combined ranking and select the top-1 predicted mimicked artist. Experiments conducted on the MIMICz dataset achieve 68% top-1 accuracy, outperforming early fusion and baseline methods, while preserving the distinct representational contributions of both spectral and prosodic feature sets. To the best of our knowledge, this is the first work to apply score level fusion of augmented spectral and prosodic embeddings for speech mimicry recognition, combining the strengths of both feature domains without compromising their individual contributions.