Point Clouds with Geometric and Semantic Confidence Intervals
摘要
To improve the application of existing natural language advancements to 3D representations, we present a novel approach for the efficient construction of 3D point clouds from image and text inputs. Unlike existing methods that focus on generating plausible views, our method extends functionality to yield accurate 3D representations much more quickly. Leveraging a CLIP-based segmentation model, our method can reconstruct the 3D representation of given data accurately, even with a limited number of input images. Furthermore, in many disciplines, reconstructing point clouds from image and text inputs will require robustness against perturbations of image and text inputs. We achieve this by encoding geometric and semantic confidence intervals for our outputs. Notably, our inclusion of semantic confidence values proves effective in mitigating false positives (i.e., identifying objects irrelevant to the text prompt), a prevalent problem among adjacent works. We believe that our work will improve the robustness of predictions and reconstructions in fields such as robotics, architectural imaging, and medical imaging.