Unified face attack detection via multi-modal multi-scale CNN–ViT network: enhancing representation capability
摘要
Face attack detection plays a vital role in safeguarding face recognition (FR) systems. Due to the unrestricted access to enormous face images and face manipulation tools circulating on the Internet, both physical and digital attacks pose significant threats to the widespread use of FR systems. However, previous works consider the detection of physical attacks and digital attacks as two independent tasks, causing an inferior generalization of attack detection among different categories. In this paper, we propose a unified multi-modal multi-scale hybrid CNN and transformer detection framework, namely