With the rapid advancements in face recognition (FR) technology, current systems perform well in unconstrained scenarios. However, detecting face spoofing attacks remains a significant challenge, making face anti-spoofing (FAS) a critical research area. Although numerous anti-spoofing models have been developed, their generalization to unseen attacks often weakens when faced with challenging variations like background, lighting, diverse spoof mediums, and low image resolution. To overcome these limitations, we propose a novel bi-branch FAS framework that leverages a pre-trained Vision Transformer (ViT) with RGB and depth data as input. The ViT's self-attention mechanism excels at capturing intricate image contexts, making it particularly effective for Presentation Attack Detection (PAD) tasks. To enhance computational efficiency, we introduce a parameter-sharing technique within the dual-branch ViT network, substantially reducing the computational burden while maintaining robust feature learning. Our framework stands out on the CASIA-FASD, Replay-Attack, and OULU-NPU benchmarks in both intra- and cross-dataset testing, while also enhancing computational efficiency.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Leveraging Depth Data and Parameter Sharing in Vision Transformers for Improved Face Anti-spoofing

  • Aashania Antil,
  • Chhavi Dhiman

摘要

With the rapid advancements in face recognition (FR) technology, current systems perform well in unconstrained scenarios. However, detecting face spoofing attacks remains a significant challenge, making face anti-spoofing (FAS) a critical research area. Although numerous anti-spoofing models have been developed, their generalization to unseen attacks often weakens when faced with challenging variations like background, lighting, diverse spoof mediums, and low image resolution. To overcome these limitations, we propose a novel bi-branch FAS framework that leverages a pre-trained Vision Transformer (ViT) with RGB and depth data as input. The ViT's self-attention mechanism excels at capturing intricate image contexts, making it particularly effective for Presentation Attack Detection (PAD) tasks. To enhance computational efficiency, we introduce a parameter-sharing technique within the dual-branch ViT network, substantially reducing the computational burden while maintaining robust feature learning. Our framework stands out on the CASIA-FASD, Replay-Attack, and OULU-NPU benchmarks in both intra- and cross-dataset testing, while also enhancing computational efficiency.