In recent years, significant progress has been made in developing multi-modal learning methods to improve face anti-spoofing systems. However, in real-world multi-modal face data, modalities from different imaging sensors are often missing. Previous studies have tended to ignore this problem, and few attempts to address this issue have made the models more complex. This study introduces a simple but robust methodology that uses a multi-modal face anti-spoofing architecture with a spatial-temporal encoders and a unit involved in fusion. The encoders of spatial-temporal extract features from each modality using ResNet34 and Transformer architectures. Augmentation and regularization techniques are used to enhance model performance. Fusion methods are evaluated for their effectiveness in handling missing modalities. Additionally, a modular autoencoder called FaceMAE is introduced to predict missing modalities. Experimental results demonstrate the robustness of FaceMAE FAS in real-world scenarios. The proposed approach is evaluated on various datasets CASIA-SURF, CASIA-SURF CeFA, and WMCA shows competitive results. Code is available https://github.com/ZainUlAbideenMalik/FaceMAE-FAS .

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Robust Multi-modal Face Anti-spoofing: Techniques for Missing Modality Handling and Fusion

  • Zain Ul Abideen,
  • Shu Liu,
  • Tongming Wan

摘要

In recent years, significant progress has been made in developing multi-modal learning methods to improve face anti-spoofing systems. However, in real-world multi-modal face data, modalities from different imaging sensors are often missing. Previous studies have tended to ignore this problem, and few attempts to address this issue have made the models more complex. This study introduces a simple but robust methodology that uses a multi-modal face anti-spoofing architecture with a spatial-temporal encoders and a unit involved in fusion. The encoders of spatial-temporal extract features from each modality using ResNet34 and Transformer architectures. Augmentation and regularization techniques are used to enhance model performance. Fusion methods are evaluated for their effectiveness in handling missing modalities. Additionally, a modular autoencoder called FaceMAE is introduced to predict missing modalities. Experimental results demonstrate the robustness of FaceMAE FAS in real-world scenarios. The proposed approach is evaluated on various datasets CASIA-SURF, CASIA-SURF CeFA, and WMCA shows competitive results. Code is available https://github.com/ZainUlAbideenMalik/FaceMAE-FAS .