<p>Monocular 3D face reconstruction, aiming to create a 3D face from a single image, has widespread application. Specifically, deformation-based parametric models are widely used due to their efficiency and robustness. Despite remarkable achievements in face geometry prediction, achieving high-fidelity human face reconstruction remains challenging due to the limitations of parametric methods. These limitations include difficulties in accurately extracting emotional features from complex facial expressions, leading to an emotional mismatch in the reconstructed model. Additionally, the parameters of the deformation models often lack clear semantic relevance, making them hard to interpret. This ambiguity prevents the models from fully leveraging their expressiveness. To address the above issues, we propose a face reconstruction method using cross-modal guidance to enhance the fidelity of complex expressions from real-world images. Our approach introduces an image–text multimodal expression encoder that characterizes the model’s parameters with descriptive text, improving semantic clarity and addressing expressive limitations. Additionally, we devise emotional difference losses that incorporate both perceptual and pixel-level analyses. This ensures that the reconstructed facial expressions accurately reflect the emotional content of the input image and avoid overemphasizing expression details. The emotion valence and arousal are regressed from the model’s predicted parameters to verify the consistency of the emotional information displayed in the reconstruction model and the input image. Experimental results show that our model is comparable to current advanced algorithms in terms of reconstruction error and outperforms them in emotional expression accuracy in reconstructed 3D models, achieving state-of-the-art results across various facial expression datasets. Codes will be available at:github.com/HeQ0219/CrossModal-FaceRecon.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Expression-driven monocular 3D face reconstruction based on cross-modal guidance

  • Xiaolong He,
  • Feipeng Da

摘要

Monocular 3D face reconstruction, aiming to create a 3D face from a single image, has widespread application. Specifically, deformation-based parametric models are widely used due to their efficiency and robustness. Despite remarkable achievements in face geometry prediction, achieving high-fidelity human face reconstruction remains challenging due to the limitations of parametric methods. These limitations include difficulties in accurately extracting emotional features from complex facial expressions, leading to an emotional mismatch in the reconstructed model. Additionally, the parameters of the deformation models often lack clear semantic relevance, making them hard to interpret. This ambiguity prevents the models from fully leveraging their expressiveness. To address the above issues, we propose a face reconstruction method using cross-modal guidance to enhance the fidelity of complex expressions from real-world images. Our approach introduces an image–text multimodal expression encoder that characterizes the model’s parameters with descriptive text, improving semantic clarity and addressing expressive limitations. Additionally, we devise emotional difference losses that incorporate both perceptual and pixel-level analyses. This ensures that the reconstructed facial expressions accurately reflect the emotional content of the input image and avoid overemphasizing expression details. The emotion valence and arousal are regressed from the model’s predicted parameters to verify the consistency of the emotional information displayed in the reconstruction model and the input image. Experimental results show that our model is comparable to current advanced algorithms in terms of reconstruction error and outperforms them in emotional expression accuracy in reconstructed 3D models, achieving state-of-the-art results across various facial expression datasets. Codes will be available at:github.com/HeQ0219/CrossModal-FaceRecon.