Instance-level cross-attention learning for fine-grained customizable face generation
摘要
Custom text-to-face generation aims to generate images based on user input prompt and reference image that both match the textual description and remain identity consistent. Currently, some methods optimize both scene and character generation through multitask learning, or focus only on facial features, thus neglecting the importance of separating and accurately modeling character description (e.g., “young woman wearing sunglasses”) and scenarioized description (e.g., “sunbathing at the beach”) in prompt, resulting in the model is not accurate enough in scene restoration and character performance. To address these issues, we propose CustomFace, a novel generation framework that ensures identity consistency while realizing scene diversity. Firstly, fine-grained identity extractor is proposed to focus on the extraction of multi-scale identity features. Secondly, the instance-level cross-attention learning strategy refines the scene and character regions for instance-level modeling, effectively solving the scene–character decoupling problem. Extensive experimental results demonstrate that our method outperforms the state of the art in identity accuracy and scene restoration.