3D Model Retrieval with Large Multimodal Models
摘要
In the field of industrial design, efficiently retrieving and utilizing existing 3D assets has always been a critical challenge. With the rapid development of multimodal large model technology, it has become feasible to achieve unified representations of text, images, and 3D models at different levels. However, existing 3D asset databases often lack intelligent retrieval mechanisms when faced with large volumes of data, making it difficult to achieve precise matching through natural language or image inputs. To address this issue, this paper proposes a 3D model retrieval method based on multimodal large models, which can quickly locate target 3D models through natural language descriptions or image inputs. This method leverages multimodal large models for feature extraction, analysis, updating, and validation without requiring specialized training on specific data. Additionally, vector retrieval technology is employed as an auxiliary means to further enhance the efficiency and effectiveness of the retrieval process. Furthermore, this paper introduces a method for constructing a 3D model database that can automatically organize 3D assets and generate an efficiently query able database. We constructed a multimodal dataset comprising images, text, and 3D models and used the proposed method to establish a 3D asset database. The retrieval performance of 3D models was tested using both textual and image inputs. Experimental results demonstrate that the proposed method excels in 3D model retrieval tasks, achieving a Top-1 accuracy rate of 91.6% for image retrieval, thereby fully validating the effectiveness of our approach.