3D reconstruction is one of the novel and meaningful tasks in aircraft vision processing. Reconstructing 3D models from 2D photos or videos taken by cameras carried on aircraft can effectively help construct ground targets, thus promoting the study of ground targets and helping researchers conduct deeper and more comprehensive work. Most of the currently validated 3D reconstruction datasets are based on indoor scenes and small objects. There has been relative little validation of outdoor building datasets captured by camera-equipped drones. Therefore, in this paper, we explore a 3D reconstruction algorithm that can be applied to scenes captured by unmanned aerial vehicles. To improve accuracy and robustness, we design an unsupervised multi-view stereo algorithm based on contrastive learning and knowledge distillation, which can predict the depth of scenes captured by drones without training from scratch. Specifically, first we use group-level contrastive loss, image-level contrastive loss and photometric consistency to build an unsupervised MVS model and train a teacher model; then we apply knowledge distillation to extract the effective probability distribution in the teacher model to train the student model, which can reduce the model size while improving accuracy and completeness; finally, the model trained on the DTU dataset is used to perform zero-shot inference on the UrbanScene dataset which contains real scenes captured by drones. It is verified that the proposed model can obtain good quantitative results on DTU and can make accurate and complete 3D reconstruction models on UrbanScene.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Unsupervised Multi-view Stereo for UAV Perspective Using Contrastive Learning and Knowledge Distillation

  • Yuhan Lu,
  • Tongtong Zhang,
  • Xuyang Wang,
  • Yuanxiang Li

摘要

3D reconstruction is one of the novel and meaningful tasks in aircraft vision processing. Reconstructing 3D models from 2D photos or videos taken by cameras carried on aircraft can effectively help construct ground targets, thus promoting the study of ground targets and helping researchers conduct deeper and more comprehensive work. Most of the currently validated 3D reconstruction datasets are based on indoor scenes and small objects. There has been relative little validation of outdoor building datasets captured by camera-equipped drones. Therefore, in this paper, we explore a 3D reconstruction algorithm that can be applied to scenes captured by unmanned aerial vehicles. To improve accuracy and robustness, we design an unsupervised multi-view stereo algorithm based on contrastive learning and knowledge distillation, which can predict the depth of scenes captured by drones without training from scratch. Specifically, first we use group-level contrastive loss, image-level contrastive loss and photometric consistency to build an unsupervised MVS model and train a teacher model; then we apply knowledge distillation to extract the effective probability distribution in the teacher model to train the student model, which can reduce the model size while improving accuracy and completeness; finally, the model trained on the DTU dataset is used to perform zero-shot inference on the UrbanScene dataset which contains real scenes captured by drones. It is verified that the proposed model can obtain good quantitative results on DTU and can make accurate and complete 3D reconstruction models on UrbanScene.