Semi-supervised Speaker Localization with Gaussian-Like Pseudo-labeling
摘要
Speaker localization is important for many human-robot interaction applications. Most existing localization studies train models with fully-annotated data through supervised learning strategies, which doesn’t fit the real-world scenarios where labeled data is scarce. To address the challenge of limited labeled data, we propose a semi-supervised deep learning algorithm for Direction of Arrival (DoA) estimation. Specifically, the model is enhanced through a process where it generates pseudo-labels for unlabeled data and incorporates a sophisticated filtering mechanism. This refined approach retrains the model by integrating both the labeled data and the data enriched with these pseudo-labels, thereby optimizing its learning capabilities. Experimental results show that it outperforms both supervised learning algorithms and other semi-supervised learning algorithms.