<p>Optical flow, which computes the apparent motion between a pair of video frames, is a critical tool for scene motion estimation. Teed-Deng’s Recurrent All-Pairs Field Transforms (RAFT) neural model based on Gated Recurrent Units (GRU), has established a new paradigm for real-time flow computation. However, we find that its correlation volume, which is the core of the network, lacks depth cues. This leads to difficulties in estimating large displacements caused by motion parallax. To address this limitation without compromising RAFT’s fast speed, this paper investigates how to reuse the depth cues latent in the network to construct a parallax-robust correlation volume. By disseminating the layers of RAFT-style networks, we find that the vectors used to initialize the GRUs’ hidden state can learn to capture depth layering. Based on our analysis, we extend the generator of the initial hidden state and utilize its depth cues to guide correlation volume construction. By integrating the proposed low-cost Parallax-Robust Correlation Volume (PRCV) with RAFT, we improve RAFT’s zero-shot flow estimation accuracy by 11.88% on Sintel Clean, and improve its fine-tuned flow estimation accuracy by 9.5% on KITTI-15, without compromising the computation efficiency. We show that PRCV also benefits other RAFT-based fast models. Comprehensive experiments validate the superiority of PRCV over the traditional correlation volume, particularly at handling motion parallax.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Parallax-robust correlation volume for optical flow computation neural networks

  • Jiang-Peng Li,
  • Yan Niu

摘要

Optical flow, which computes the apparent motion between a pair of video frames, is a critical tool for scene motion estimation. Teed-Deng’s Recurrent All-Pairs Field Transforms (RAFT) neural model based on Gated Recurrent Units (GRU), has established a new paradigm for real-time flow computation. However, we find that its correlation volume, which is the core of the network, lacks depth cues. This leads to difficulties in estimating large displacements caused by motion parallax. To address this limitation without compromising RAFT’s fast speed, this paper investigates how to reuse the depth cues latent in the network to construct a parallax-robust correlation volume. By disseminating the layers of RAFT-style networks, we find that the vectors used to initialize the GRUs’ hidden state can learn to capture depth layering. Based on our analysis, we extend the generator of the initial hidden state and utilize its depth cues to guide correlation volume construction. By integrating the proposed low-cost Parallax-Robust Correlation Volume (PRCV) with RAFT, we improve RAFT’s zero-shot flow estimation accuracy by 11.88% on Sintel Clean, and improve its fine-tuned flow estimation accuracy by 9.5% on KITTI-15, without compromising the computation efficiency. We show that PRCV also benefits other RAFT-based fast models. Comprehensive experiments validate the superiority of PRCV over the traditional correlation volume, particularly at handling motion parallax.