IAFF-VC: Any-to-Any Voice Conversion Using Attentional Feature Fusion
摘要
Voice conversion (VC) requires precise preservation of source speech content while effectively capturing target speaker characteristics. Current approaches predominantly focus on extracting granular acoustic representations, yet often neglect the crucial integration of heterogeneous feature types, resulting in an inherent performance trade-off between content fidelity and speaker similarity. To address this limitation, we present IAFF-VC, a novel framework for non-parallel any-to-any voice conversion that synergistically combines encoder–decoder architecture with multi-scale feature fusion. Our proposed EMAFF module introduces a multi-branch channel grouping mechanism that strategically reorganizes spatial-semantic features across distinct subspaces, enabling optimal fusion of complementary speech attributes. Comprehensive evaluations on the VCTK benchmark demonstrate IAFF-VC’s superior performance in one-shot conversion scenarios, achieving state-of-the-art results with 9.75% CER and 91.76% speaker similarity score, while maintaining 4.21 mean opinion score for speech naturalness.