An intermediate frame is generated by video frame interpolation among two original sequential video frames. Vehicle driving videos contain a large number of complex scenes with large movements and occlusions. Existing frame interpolation methods based on convolutional kernel estimation and optical flow estimation have achieved favorable results when dealing with static or small-scale motion scenes. However, the limitations of the local receptive fields of the convolutional kernel and the range of optical flow estimation make these methods perform poorly when dealing with large motions. To address this problem, we provide the vehicle video frame interpolation based on spatial-channel reconstruction and transformers (SCR-T), designed to handle video frame interpolation in complex scenes with large motions. Specifically, SCR-T initially uses the spatial-channel reconstruction module (ScConv) to extract motion and background information efficiently. Subsequently, employing two consecutive Transformer block structures and introducing relative position bias to capture long distance dependencies, this process generates appearance and motion information. Finally, the intermediate optical flow is estimated based on the feature information used to generate an intermediate frame. Our method and previous methods are tested on the datasets of Driving Video with Object Tracking and Driving videos on YouTube respectively. Through extensive quantitative and qualitative experiments, it has been shown that our method achieves better results on these datasets.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

SCR-T: Vehicle Video Frame Interpolation Based on Spatial-Channel Reconstruction and Transformers

  • Weiwei Xie,
  • Jin Lv,
  • Xulong Li

摘要

An intermediate frame is generated by video frame interpolation among two original sequential video frames. Vehicle driving videos contain a large number of complex scenes with large movements and occlusions. Existing frame interpolation methods based on convolutional kernel estimation and optical flow estimation have achieved favorable results when dealing with static or small-scale motion scenes. However, the limitations of the local receptive fields of the convolutional kernel and the range of optical flow estimation make these methods perform poorly when dealing with large motions. To address this problem, we provide the vehicle video frame interpolation based on spatial-channel reconstruction and transformers (SCR-T), designed to handle video frame interpolation in complex scenes with large motions. Specifically, SCR-T initially uses the spatial-channel reconstruction module (ScConv) to extract motion and background information efficiently. Subsequently, employing two consecutive Transformer block structures and introducing relative position bias to capture long distance dependencies, this process generates appearance and motion information. Finally, the intermediate optical flow is estimated based on the feature information used to generate an intermediate frame. Our method and previous methods are tested on the datasets of Driving Video with Object Tracking and Driving videos on YouTube respectively. Through extensive quantitative and qualitative experiments, it has been shown that our method achieves better results on these datasets.