Novel views synthesis from sparse input views is a long-standing and critical task in the field of 3D reconstruction. Existing representative methods that advance this field are mainly based on Neural Radiance Fields (NeRF), which are limited by their per-scene lengthy optimization time and high computational overhead for rendering. Recent advances on 3D Gaussian Splatting (3D-GS)  technique have shown impressive potential, enabling fast training in 10 min and real time rendering by up to \(50 \times \) speedup without sacrificing quality. However, in sparse view settings, 3D-GS overfits the training views, resulting in significant performance drops on testing views and severe floater artifacts visually. We observe that this performance degradation is due to the geometry and appearance are extremely under-constrained in a few-shot setting. To address this issue and prevent overfitting, we utilize off-the-shelf models to guide the optimization of sparse input in 3D-GS. Our key idea is to combine two types of complementary geometry priors to supervise 3D-GS training. Specifically, we leverage the depth from a monocular depth estimation network and the image correspondence distances from a geometric matching network as priors. With regularization, 3D-GS learns more geometrically consistent structure, has generalization capabilities, and enables realistic novel view synthesis. With regularization, 3D-GS learns more geometrically consistent structures, gains generalization capabilities, and can render realistic novel view. Experiments on various real-world datasets show that our straightforward yet effective approach surpasses previous methods. Code is available at: https://github.com/leo-frank/diff-gaussian-rasterization-depth

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Regularizing Sparse Input 3D Gaussian Splatting with Geometry Priors

  • Wenyu Li,
  • Ziteng Zhang,
  • Zongxin Ye,
  • Sidun Liu,
  • Peng Qiao

摘要

Novel views synthesis from sparse input views is a long-standing and critical task in the field of 3D reconstruction. Existing representative methods that advance this field are mainly based on Neural Radiance Fields (NeRF), which are limited by their per-scene lengthy optimization time and high computational overhead for rendering. Recent advances on 3D Gaussian Splatting (3D-GS)  technique have shown impressive potential, enabling fast training in 10 min and real time rendering by up to \(50 \times \) speedup without sacrificing quality. However, in sparse view settings, 3D-GS overfits the training views, resulting in significant performance drops on testing views and severe floater artifacts visually. We observe that this performance degradation is due to the geometry and appearance are extremely under-constrained in a few-shot setting. To address this issue and prevent overfitting, we utilize off-the-shelf models to guide the optimization of sparse input in 3D-GS. Our key idea is to combine two types of complementary geometry priors to supervise 3D-GS training. Specifically, we leverage the depth from a monocular depth estimation network and the image correspondence distances from a geometric matching network as priors. With regularization, 3D-GS learns more geometrically consistent structure, has generalization capabilities, and enables realistic novel view synthesis. With regularization, 3D-GS learns more geometrically consistent structures, gains generalization capabilities, and can render realistic novel view. Experiments on various real-world datasets show that our straightforward yet effective approach surpasses previous methods. Code is available at: https://github.com/leo-frank/diff-gaussian-rasterization-depth