Cross-Media Correlation Computational Method for Multimodal Semantic Sparse Data
摘要
For the cross-media retrieval, due to the “heterogeneous gap” problem between different modal data caused by the existence of the characterization of inconsistencies, it is difficult to measure the similarity directly. To solve this problem, we propose a Cross-media Correlation Computational Method for Multimodal Semantic Sparse Data (CCCM). CCCM supports five modal semantic sparse data, including text, image, video, audio, and 3D model. First, we propose a fine-grained cross-media correlation learning method that fuses cross-entropy and distribution differences. The data in the dataset is used for cross-media correlation learning through fine-grained segmentation, LSTM network, and loss function that combines cross-entropy and distribution difference. On this basis, most existing methods only consider the correlation analysis between different media data and ignore the sparse semantic problem existing in cross-media datasets. So we proposed a keyword-based KTF-IDF method to quantify the correlation between semantic tags. By conducting experiments on a large-scale cross-media dataset containing five modal data, compared with the existing common methods, the experimental results show that the correlation obtained by CCCM can be applied to multi-modal data cross-media. It can improve the accuracy by an average of 23% when it comes to media retrieval tasks.