Thermal imaging technology is crucial for biometric recognition and security monitoring, as it offers a reliable means of identification across various lighting conditions. However, despite the reliance of most facial verification systems on thermal images, obtaining pairs of visible-thermal images is often impractical due to high costs, stringent legal policies, and environmental factors. To address this challenge, we propose GTF, a framework for generating identity-preserving visible-thermal face image pairs from single-modality inputs. Our approach combines two key innovations: (1) stacked ID embedding for multi-image identity consolidation, and (2) hybrid-attention DDPM for spectral-aware translation. GTF first employs Long-Photomaker to generate diverse visible faces with consistent identity, then transforms them into thermal images via an attention-gated diffusion process that preserves facial biometric features. Experiments on SpeakingFaces and ARL_VTF datasets demonstrate that GTF establishes the new state-of-the-art in facial thermal generation. Specifically, it outperforms existing methods in both stages: (i) initial visible face generation (VR@FAR = 75.67 on SpeakingFaces) and (ii) thermal translation (SSIM = 79.05), achieving end-to-end superiority in cross-modal fidelity. This work enables privacy-compliant synthetic data generation for security applications while guaranteeing both pairwise modality consistency and identity retention.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

GTF: Generator for Pairwise Thermal Face Image Synthesis

  • Jiankai Lu,
  • Cheng Zhao,
  • Ying Zhang,
  • Jinyu Zhu,
  • Jialei Zheng

摘要

Thermal imaging technology is crucial for biometric recognition and security monitoring, as it offers a reliable means of identification across various lighting conditions. However, despite the reliance of most facial verification systems on thermal images, obtaining pairs of visible-thermal images is often impractical due to high costs, stringent legal policies, and environmental factors. To address this challenge, we propose GTF, a framework for generating identity-preserving visible-thermal face image pairs from single-modality inputs. Our approach combines two key innovations: (1) stacked ID embedding for multi-image identity consolidation, and (2) hybrid-attention DDPM for spectral-aware translation. GTF first employs Long-Photomaker to generate diverse visible faces with consistent identity, then transforms them into thermal images via an attention-gated diffusion process that preserves facial biometric features. Experiments on SpeakingFaces and ARL_VTF datasets demonstrate that GTF establishes the new state-of-the-art in facial thermal generation. Specifically, it outperforms existing methods in both stages: (i) initial visible face generation (VR@FAR = 75.67 on SpeakingFaces) and (ii) thermal translation (SSIM = 79.05), achieving end-to-end superiority in cross-modal fidelity. This work enables privacy-compliant synthetic data generation for security applications while guaranteeing both pairwise modality consistency and identity retention.