FedSER-XAI: PSO-optimized multi-stream cross-attention transformer with graph features for explainable federated speech emotion recognition
摘要
Federated learning for speech emotion recognition faces fundamental challenges in simultaneously achieving high performance, privacy preservation, and model interpretability. This paper introduces FedSER-XAI, a novel framework that integrates Particle Swarm Optimization (PSO)-based feature selection, multi-stream cross-attention mechanisms, and graph-based feature extraction within an explainable federated learning architecture. Our approach combines Vision Transformer processing of mel-spectrograms with temporal-spatial graph convolutional networks to capture both contextual and structural speech relationships. The PSO algorithm achieves 78.1% dimensionality reduction (228