GTF: Generator for Pairwise Thermal Face Image Synthesis
摘要
Thermal imaging technology is crucial for biometric recognition and security monitoring, as it offers a reliable means of identification across various lighting conditions. However, despite the reliance of most facial verification systems on thermal images, obtaining pairs of visible-thermal images is often impractical due to high costs, stringent legal policies, and environmental factors. To address this challenge, we propose GTF, a framework for generating identity-preserving visible-thermal face image pairs from single-modality inputs. Our approach combines two key innovations: (1) stacked ID embedding for multi-image identity consolidation, and (2) hybrid-attention DDPM for spectral-aware translation. GTF first employs Long-Photomaker to generate diverse visible faces with consistent identity, then transforms them into thermal images via an attention-gated diffusion process that preserves facial biometric features. Experiments on SpeakingFaces and ARL_VTF datasets demonstrate that GTF establishes the new state-of-the-art in facial thermal generation. Specifically, it outperforms existing methods in both stages: (i) initial visible face generation (VR@FAR = 75.67 on SpeakingFaces) and (ii) thermal translation (SSIM = 79.05), achieving end-to-end superiority in cross-modal fidelity. This work enables privacy-compliant synthetic data generation for security applications while guaranteeing both pairwise modality consistency and identity retention.