<p>Accurate prediction of human skeletal motion sequences is critical for human activity analysis and low-latency motion reconstruction applications. While many studies focus on frame-by-frame prediction model designs, the keyframes in a motion sequence may contain more spatial-temporal information than the other keyframes do. To address the importance of keyframes, this work introduces a heterogeneous keyframe selection and fusion method to discriminate the importance of different motion frames from historical observations for prediction. Specifically, we propose an adaptive keyframe selection algorithm to iteratively select the keyframes and a nonlinear heterogeneous interpolation method to reconstruct the transitional frames. By merging them with the original motion sequence, the semantics of the original motion are preserved, and the importance of the keyframes is highlighted. A graph convolutional network (GCN) is designed for prediction with dual-channel attention to incorporate motion patterns in longer-term historical records to improve motion feature exploration. A comprehensive evaluation of the model is performed on the Human3.6M and AMASS datasets, which shows significant improvement in motion prediction over long-term methods (<InlineEquation ID="IEq1"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="10489_2025_6532_Article_IEq1.gif" Format="GIF" Height="15" Rendition="HTML" Resolution="72" Type="Linedraw" Width="19" /> </InlineMediaObject> <EquationSource Format="TEX">\(\ge \)</EquationSource> <EquationSource Format="MATHML"><math> <mo>≥</mo> </math></EquationSource> </InlineEquation> 320 ms) over the state-of-the-art methods in terms of the 3D mean per joint position error (MPJPE).</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A keyframe weighted dual-channel attention GCN model for human skeleton motion prediction

  • Wenwen Zhang,
  • Jianfeng Tu,
  • Siyu Li,
  • Lingfeng Liu

摘要

Accurate prediction of human skeletal motion sequences is critical for human activity analysis and low-latency motion reconstruction applications. While many studies focus on frame-by-frame prediction model designs, the keyframes in a motion sequence may contain more spatial-temporal information than the other keyframes do. To address the importance of keyframes, this work introduces a heterogeneous keyframe selection and fusion method to discriminate the importance of different motion frames from historical observations for prediction. Specifically, we propose an adaptive keyframe selection algorithm to iteratively select the keyframes and a nonlinear heterogeneous interpolation method to reconstruct the transitional frames. By merging them with the original motion sequence, the semantics of the original motion are preserved, and the importance of the keyframes is highlighted. A graph convolutional network (GCN) is designed for prediction with dual-channel attention to incorporate motion patterns in longer-term historical records to improve motion feature exploration. A comprehensive evaluation of the model is performed on the Human3.6M and AMASS datasets, which shows significant improvement in motion prediction over long-term methods ( \(\ge \) 320 ms) over the state-of-the-art methods in terms of the 3D mean per joint position error (MPJPE).