<p>Recent researches on deep-learning methods have shown enormous promise for depth estimates. Although self-supervised estimators alleviate laborious annotation, numerous pixels fail to match the corresponding ones in adjacent frames. Single re-projection loss functions are oversimplified for stereoscopic adjacent frame constraints. In this work, we constrain the primitive inter-frame-supervised depth estimation via multiple bilateral consistency, which builds a bi-directional mapping between adjacent frames with inherent properties, allowing to develop pose-consistent and depth-consistent models. For the mismatching pixels in adjacent frames, a cycle-consistency framework reconstructs depth maps as scene images with an additional symmetrical re-rendering network. Pose-consistent constraint aims to ensure the reversibility of ego-motion transformations between adjacent frames, and depth-consistent constraint strives to ensure the continuity of adjacent frames’ depths. With the joint optimization of three independent modules, depth network, pose network and re-rendering network, the proposed framework yields state-of-the-art metrics on the KITTI depth dataset pre-trained on CityScapes, with the AbsRel, SqRel, RMSE and RMS(log) decreased by 6.6%, 12.4%, 0.4%, and 2.2%, respectively.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Self-supervised monocular depth estimation via multiple bilateral consistency

  • Zhengyang Lu,
  • Ying Chen

摘要

Recent researches on deep-learning methods have shown enormous promise for depth estimates. Although self-supervised estimators alleviate laborious annotation, numerous pixels fail to match the corresponding ones in adjacent frames. Single re-projection loss functions are oversimplified for stereoscopic adjacent frame constraints. In this work, we constrain the primitive inter-frame-supervised depth estimation via multiple bilateral consistency, which builds a bi-directional mapping between adjacent frames with inherent properties, allowing to develop pose-consistent and depth-consistent models. For the mismatching pixels in adjacent frames, a cycle-consistency framework reconstructs depth maps as scene images with an additional symmetrical re-rendering network. Pose-consistent constraint aims to ensure the reversibility of ego-motion transformations between adjacent frames, and depth-consistent constraint strives to ensure the continuity of adjacent frames’ depths. With the joint optimization of three independent modules, depth network, pose network and re-rendering network, the proposed framework yields state-of-the-art metrics on the KITTI depth dataset pre-trained on CityScapes, with the AbsRel, SqRel, RMSE and RMS(log) decreased by 6.6%, 12.4%, 0.4%, and 2.2%, respectively.