D-NeuRA: customizable dynamic neural rendering for human avatars with disentangled pose and appearance
摘要
Neural Radiance Fields (NeRF) has become a fundamental method for generating human avatars due to its successful application in novel view synthesis. Existing NeRF-based methods build static features that are independent of pose to learn and represent human albedo and shadow information, combined with external lighting information to generate fully customizable human avatars. However, these static features can only provide global, pose-independent information and are unable to adapt to the detailed variations caused by pose and lighting changes. To address this, we propose a human avatar modeling framework, D-NeuRA, which integrates static theme priors and image-driven dynamic features. Specifically, we map three-dimensional sampling points to image coordinates and perform bilinear sampling on image feature maps to extract dynamic human appearance features. We also design a Multi-Scale Bitemporal Dynamic Feature Fusion module to effectively combine dynamic features with static theme features to generate adaptive appearance features. To further enhance the expression of pose modeling at the local geometric level, we introduce a two-stage pose feature extraction and enhancement strategy. In the first stage, key pose retrieval and interpolation are used to sample and extract preliminary pose features along three axes. In the second stage, a Median-Enhanced Channel and Spatial Attention (MECS) module refines and enhances the pose features. Finally, we use multiple decoders to combine the adaptive appearance features and enhanced pose features to generate neural field properties, rendering high-quality, fully customizable dynamic human avatars. Extensive experiments show that the proposed method outperforms existing mainstream approaches in rendering realism, pose consistency, and detail preservation, particularly demonstrating stronger robustness and generalization ability in complex dynamic scenes.