DictAvatar: Expressive Facial Avatar Reconstruction with Facial Feature Dictionary
摘要
Existing NeRF-based head avatar reconstruction methods utilize expression coefficients as driving signals. Despite significant advancements, they fail to accurately capture facial feature deformations under complex expression changes. To address this issue, we integrate prior information from public facial feature dictionaries with expression coefficients as driving signals, and employ a region attention mechanism to more accurately capture facial deformations. First, when reconstructing a single face, we extract the facial features from a single image and obtain prior information on these features from a public facial feature dictionary. This prior information is integrated with expression coefficients as driving signals to more accurately drive the deformation of facial features. Second, we introduce a region attention mechanism that learns the explicit relationship between local spatial regions and driving signals during training. This allows for differential driving effects on various facial regions with the same driving signal, achieving more precise local motion modeling. Experimental results show that our method can precisely capture subtle deformations across all facial regions and outperforms state-of-the-art methods in both qualitative and quantitative aspects.