Attention-Guided Contrastive Masked Autoencoders for Self-supervised Cross-Modal Biometric Matching
摘要
Cross-modal biometric matching is a significant research area in artificial intelligence and information forensics, focusing on establishing correlation between human acoustic and facial features. However, the variation of viewpoints in different scenes can lead to incomplete representations of specific voice and face images, which cannot be used to learn local critical features in the samples in an insightful way. To address these issues, we developed the Attention-Guided Contrast Masking Autoencoder Framework (ACMAE). This framework enhances cross-modal biometric performance through contrastive learning and masked data reconstruction. First, we introduce an attention-based data reconstruction method to improve feature discrimination. Second, we design multi-view contrastive learning to reduce viewpoint differences and enhance the robustness of the representations. Finally, experiments on benchmark datasets demonstrate that the ACMAE method achieves state-of-the-art performance and proves its effectiveness.