Multi-view Depth Estimation with Adaptive Feature Extraction and Region-Aware Depth Prediction
摘要
Multi-view depth estimation is an essential task for 3D reconstruction, which aims to obtain depth map from multi-view images via multi-view stereo (MVS) technique. However, when the input images contain challenging regions with occlusion and low texture, existing MVS methods may fail to perform well. To tackle these problems, this paper proposes a multi-view depth estimation framework, which consists of adaptive feature extraction and region-aware depth prediction. To obtain better pixel feature matching, adaptive feature extraction is constructed with a CNN-based Adaptive Local Feature Extractor (ALFE) and a Transformer-based Global Feature Extractor (GFE) to capture the representative and robust features in challenging regions. To obtain better depth map output, region-aware depth prediction is constructed with the Region-Aware Depth Refinement Module (RA-DRM), which iteratively refines the depth map guided by the extracted features in different regions. Qualitative and quantitative experiments were conducted on the DTU and BlendedMVS datasets. The results show the effectiveness of the proposed modules of ALFE, GFE and RA-DRM, and the comparisons indicate that our learning-based MVS method is superior to related methods.