<p>Robotic motor control necessitates the ability to predict the dynamics of interactions with environment and objects. However, existing self-supervised pre-trained visual representations in robotic motor control, leveraging large-scale egocentric videos, often focus solely on learning the static content features. Thus it neglects the crucial motion clues in human videos, which contain key knowledge about interaction with the environments and objects. In this paper, we present a simple yet effective visual pre-training framework termed as <b>STP</b> for robotic motor control, which jointly performs spatial and temporal prediction with different masking ratios and a dual-decoder design by utilizing large-scale egocentric video data. Specifically, given paired frames in a video clip, we first employ a spatial decoder to perform spatial prediction on the masked current frame to learn the content features. Second, we employ a separate temporal decoder to conduct temporal prediction, conditioned on the masked current frame and an extremely highly-masked future frame, for capturing motion features. This <b>asymmetric masking and decoupled decoding design</b> differentiate our STP from the traditional image or video masked autoencoders by explicitly ensuring the <b>learned image representation to focus more on motion information while still capture spatial details</b>. Due to its lightweight temporal decoder and highly masked future frame, STP incurs manageable extra training overhead over standard image MAE. Extensive simulation and real-world experiments demonstrate the effectiveness and generalization abilities of STP, especially in generalizing to unseen environments with more distractors. Additionally, we extend this framework with post-training and hybrid pre-training to demonstrate the generality and data efficiency of STP. The source code and models will be available at <a href="https://github.com/MCG-NJU/STP.">https://github.com/MCG-NJU/STP.</a></p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Spatiotemporal Predictive Pre-training for Robotic Motor Control

  • Jiange Yang,
  • Bei Liu,
  • Jianlong Fu,
  • Bocheng Pan,
  • Gangshan Wu,
  • Limin Wang

摘要

Robotic motor control necessitates the ability to predict the dynamics of interactions with environment and objects. However, existing self-supervised pre-trained visual representations in robotic motor control, leveraging large-scale egocentric videos, often focus solely on learning the static content features. Thus it neglects the crucial motion clues in human videos, which contain key knowledge about interaction with the environments and objects. In this paper, we present a simple yet effective visual pre-training framework termed as STP for robotic motor control, which jointly performs spatial and temporal prediction with different masking ratios and a dual-decoder design by utilizing large-scale egocentric video data. Specifically, given paired frames in a video clip, we first employ a spatial decoder to perform spatial prediction on the masked current frame to learn the content features. Second, we employ a separate temporal decoder to conduct temporal prediction, conditioned on the masked current frame and an extremely highly-masked future frame, for capturing motion features. This asymmetric masking and decoupled decoding design differentiate our STP from the traditional image or video masked autoencoders by explicitly ensuring the learned image representation to focus more on motion information while still capture spatial details. Due to its lightweight temporal decoder and highly masked future frame, STP incurs manageable extra training overhead over standard image MAE. Extensive simulation and real-world experiments demonstrate the effectiveness and generalization abilities of STP, especially in generalizing to unseen environments with more distractors. Additionally, we extend this framework with post-training and hybrid pre-training to demonstrate the generality and data efficiency of STP. The source code and models will be available at https://github.com/MCG-NJU/STP.