<p>The limitations of single-modality control in preserving facial identity and describing temporal expressions have motivated sketch-text dual-driven 4D face generation, which provides a flexible paradigm for dynamic digital face synthesis in applications requiring precise identity customization and controllable expression manipulation. However, this task remains challenging due to the synthetic-to-real domain gap in sparse sketches, cross-modal interference between heterogeneous conditions, and the scarcity of paired sketch-text 4D mesh data. To address these challenges, we propose sketch-text-driven rectified flow (STDRF), a conditional rectified-flow framework for identity-preserving and semantically controllable 4D face generation. STDRF learns a velocity field that transports Gaussian noise to the target 4D facial motion manifold under sketch-structural and text-semantic dual guidance. To obtain robust identity priors from sparse sketches, we design a sketch encoder enhanced by Geometric Contour and Texture Detail (GCTD) preprocessing and MixStyle domain adaptation. To reduce cross-modal interference, a dual-path independent cross-attention module based on IP-Adapter is designed to inject sketch features and text semantics into an Attention DiffusionNet Block (ADNB)-based denoising backbone in parallel. Furthermore, a multimodal triplet dataset is constructed by pairing 4D facial mesh sequences with synthetic sketches and hierarchical text descriptions. On the subject-independent test split, STDRF achieves an MVE of 0.81 <InlineEquation ID="IEq1"> <EquationSource Format="TEX">\(\times \)</EquationSource> <EquationSource Format="MATHML"><math> <mo>×</mo> </math></EquationSource> </InlineEquation> <InlineEquation ID="IEq2"> <EquationSource Format="TEX">\(10^{-3}\)</EquationSource> <EquationSource Format="MATHML"><math> <msup> <mn>10</mn> <mrow> <mo>-</mo> <mn>3</mn> </mrow> </msup> </math></EquationSource> </InlineEquation>, an ID-Sim of 0.989, and a GPU inference speed of 21.89 FPS for 160-frame sequences, showing favorable geometric fidelity, identity consistency, and efficiency compared with representative cascaded methods. Code is available at: <a href="https://github.com/alex11782/STDRF">https://github.com/alex11782/STDRF</a>.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Sketch-text-driven rectified flow for identity-preserving 4D face generation

  • Baodong Wang,
  • Fang Liu,
  • Wei Cao,
  • Guodong Wang,
  • Junli Zhao,
  • Zhenkuan Pan

摘要

The limitations of single-modality control in preserving facial identity and describing temporal expressions have motivated sketch-text dual-driven 4D face generation, which provides a flexible paradigm for dynamic digital face synthesis in applications requiring precise identity customization and controllable expression manipulation. However, this task remains challenging due to the synthetic-to-real domain gap in sparse sketches, cross-modal interference between heterogeneous conditions, and the scarcity of paired sketch-text 4D mesh data. To address these challenges, we propose sketch-text-driven rectified flow (STDRF), a conditional rectified-flow framework for identity-preserving and semantically controllable 4D face generation. STDRF learns a velocity field that transports Gaussian noise to the target 4D facial motion manifold under sketch-structural and text-semantic dual guidance. To obtain robust identity priors from sparse sketches, we design a sketch encoder enhanced by Geometric Contour and Texture Detail (GCTD) preprocessing and MixStyle domain adaptation. To reduce cross-modal interference, a dual-path independent cross-attention module based on IP-Adapter is designed to inject sketch features and text semantics into an Attention DiffusionNet Block (ADNB)-based denoising backbone in parallel. Furthermore, a multimodal triplet dataset is constructed by pairing 4D facial mesh sequences with synthetic sketches and hierarchical text descriptions. On the subject-independent test split, STDRF achieves an MVE of 0.81 \(\times \) × \(10^{-3}\) 10 - 3 , an ID-Sim of 0.989, and a GPU inference speed of 21.89 FPS for 160-frame sequences, showing favorable geometric fidelity, identity consistency, and efficiency compared with representative cascaded methods. Code is available at: https://github.com/alex11782/STDRF.