<p>Pose transfer refers to transferring a given person’s pose to the target pose. We present a context-driven omni-dimensional dynamic pose transfer model to address the issue of regular convolutional networks being unable to manage complicated changes. First, we construct a dynamic convolution module to extract rich contextual features. This module dynamically adjusts to the differences in input data during the convolution process, enhancing the adaptability of features. Second, a feature fusion block (FFBlock) is built by merging multiscale channel attention information from global and local channel contexts. Furthermore, the focal-frequency distance between the generated image and the original image is measured using the focal-frequency loss, which allows a model to adaptively focus on difficult-to-synthesis frequency components by reducing the weighting of easy-to-synthesis frequency components, narrowing the gap in the frequency domain, and improving image generation quality. The effectiveness and efficiency of the network are qualitatively and quantitatively verified on fashion datasets, and a large number of experiments demonstrate the superiority of our method.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

CDPT: context-driven omni-dimensional dynamic pose transfer network

  • Yue Chen,
  • Xiaoman Liang,
  • Mugang Lin,
  • Yuan Qin,
  • Huihuang Zhao

摘要

Pose transfer refers to transferring a given person’s pose to the target pose. We present a context-driven omni-dimensional dynamic pose transfer model to address the issue of regular convolutional networks being unable to manage complicated changes. First, we construct a dynamic convolution module to extract rich contextual features. This module dynamically adjusts to the differences in input data during the convolution process, enhancing the adaptability of features. Second, a feature fusion block (FFBlock) is built by merging multiscale channel attention information from global and local channel contexts. Furthermore, the focal-frequency distance between the generated image and the original image is measured using the focal-frequency loss, which allows a model to adaptively focus on difficult-to-synthesis frequency components by reducing the weighting of easy-to-synthesis frequency components, narrowing the gap in the frequency domain, and improving image generation quality. The effectiveness and efficiency of the network are qualitatively and quantitatively verified on fashion datasets, and a large number of experiments demonstrate the superiority of our method.