Indexing High-dimensional Multi-modal Big Data
摘要
Indexing high-dimensional big data with different types of data (multi-modal data) is an exciting and rapidly advancing area. Multi-modal data encompasses diverse formats such as text, images, audio, and numerical data, necessitating the use of efficient techniques that can seamlessly integrate and handle these varied data types. In this paper, we investigate the most recent techniques and approaches employed in the indexing of high-dimensional multi-modal data, including multi-modal hashing, embedding-based methods, deep metric learning, self-supervised learning, and graph-based indexing. We provide a comparative analysis of these methods across different applications, highlighting their advantages and limitations. Furthermore, we address key challenges in multi-modal data retrieval and propose future improvements. In particular, we aim to integrate K-NN with embedding-based models to enhance the retrieval performance of high-dimensional multi-modal datasets, surpassing traditional indexing approaches.