<p>Although face recognition systems are widely used, they are vulnerable to presentation attacks. Currently, there are various technical solutions for face anti-spoofing, including methods based on RGB image analysis and multimodal fusion methods based on deep learning. However, the single RGB image analysis method has insufficient generalization ability for high-quality attacks. Meanwhile, the existing multimodal fusion networks usually have high computational complexity and it is difficult to balance the feature dependencies among different modalities. The proposed method innovatively integrates RGB and depth information at different scales, simultaneously learns modality-independent features, and discovers potential deceptive features in the RGB mode. As a result, it significantly improves the generalization ability of the model. Additionally, a multi-window shuffle fusion module is designed, which establishes long-range dependency relationships while reducing the computational complexity of the Vision Transformer’s self-attention mechanism. During the training phase, a deep image-guided RGB feature learning method was adopted, while in the testing phase, only RGB data need to be input. This not only maintains the performance advantage of the multimodal system, but also meets the convenience requirements of practical applications. Extensive experiments on the CASIA-SURF, CASIA-CeFA, and WMCA datasets demonstrate that the performance of our model significantly outperforms the existing state-of-the-art methods. We achieved an ACER of 0.41% on CASIA-SURF and an ACER of 4.06% ± 1.16% on CASIA-CeFA. Through the optimized attention module, we have achieved an outstanding lightweight level, with FLOPs and parameters of only 3.52G and 8.04M, respectively.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Enhanced Depth-Guided RGB Image Feature Fusion for Robust Face Anti-spoofing

  • Tong Guoxiang,
  • Shao Haitao,
  • Yan Xinrong

摘要

Although face recognition systems are widely used, they are vulnerable to presentation attacks. Currently, there are various technical solutions for face anti-spoofing, including methods based on RGB image analysis and multimodal fusion methods based on deep learning. However, the single RGB image analysis method has insufficient generalization ability for high-quality attacks. Meanwhile, the existing multimodal fusion networks usually have high computational complexity and it is difficult to balance the feature dependencies among different modalities. The proposed method innovatively integrates RGB and depth information at different scales, simultaneously learns modality-independent features, and discovers potential deceptive features in the RGB mode. As a result, it significantly improves the generalization ability of the model. Additionally, a multi-window shuffle fusion module is designed, which establishes long-range dependency relationships while reducing the computational complexity of the Vision Transformer’s self-attention mechanism. During the training phase, a deep image-guided RGB feature learning method was adopted, while in the testing phase, only RGB data need to be input. This not only maintains the performance advantage of the multimodal system, but also meets the convenience requirements of practical applications. Extensive experiments on the CASIA-SURF, CASIA-CeFA, and WMCA datasets demonstrate that the performance of our model significantly outperforms the existing state-of-the-art methods. We achieved an ACER of 0.41% on CASIA-SURF and an ACER of 4.06% ± 1.16% on CASIA-CeFA. Through the optimized attention module, we have achieved an outstanding lightweight level, with FLOPs and parameters of only 3.52G and 8.04M, respectively.