SaLon3R: Structure-Aware Long-Term Feedforward 3D Reconstruction from Unposed Images
摘要
Recent advances in 3D Gaussian Splatting (3DGS) have enabled feed-forward, on-the-fly reconstruction of sequential input views. However, existing methods often predict per-pixel Gaussians and combine Gaussians from all views as the scene representation, leading to substantial redundancies and geometric inconsistencies in long-duration video sequences. To address this, we propose SaLon3R, a novel framework for Structure-aware, Long-term 3DGS Reconstruction. Our method eliminates redundancy by introducing compact anchor primitives as a replacement for per-pixel Gaussians. These primitives are derived through a differentiable, saliency-aware Gaussian quantization process designed to preserve fidelity while ensuring a compact representation. Specifically, a foundational 3D reconstruction model is employed to predict a saliency map encoding regional geometric complexity. Guided by this saliency map, we compress redundant Gaussian primitives into compact anchors by prioritizing high-complexity regions. Furthermore, we introduce a 3D Point Transformer to overcome geometric inconsistencies caused by long-term accumulative errors. It refines attributes and saliency of the anchor primitives leveraging the learned spatial structural priors in 3D space. Without known camera parameters or test-time optimization, our approach effectively prunes the redundant 3DGS and resolves artifacts in a single feed-forward pass. Experiments on multiple datasets demonstrate our approach outperform state-of-the-arts on both novel view synthesis and depth estimation, while exhibiting superior efficiency (