Audio-Guided Visual Knowledge Representation
摘要
Visual knowledge is primarily acquired through visual perception, but it is often exclusively represented in natural language, neglecting the collaborative nature of multisensory perception. To address this limitation, this paper proposes an audio-guided approach to visual knowledge representation. By integrating auditory cues into visual captioning, the model enhances environmental understanding through multisensory collaboration. Furthermore, the introduction of an audio-visual multimodal mutual information graph enriches the semantic content of visual captions. Additionally, while research on multimodal perception data is extensive, audio-visual datasets often lack fine-grained annotations. To address this issue, we construct a fine-grained multimodal dataset. Finally, experimental validation through multimodal-guided visual captioning and link prediction tasks demonstrates the effectiveness of this approach compared to existing methods.