CMPIR: cross-modal pose image reconstruction via style-semantic fusion
摘要
With the ubiquity of cameras as mainstream sensing infrastructure, privacy concerns have escalated. This paper introduces CMPIR, a novel cross-modal transformation framework that employs low-resolution infrared sensors to facilitate user pose image reconstruction without privacy leakage. Technically, CMPIR designs a cross-modal pose image generation model based on conditional generative adversarial networks, integrating the pose semantics of infrared heatmaps with the style attributes of RGB images to generate detailed virtual images suitable for various privacy applications. Meanwhile, to address the challenge of severe semantic information loss in infrared heatmaps, we design a novel pixel-fusion-based data augmentation algorithm, enhancing the ability to extract semantic pose information from infrared heatmaps. We collect