<p>Effective scene representation requires to encompass a complete understanding of the scene’s structure, including multimodal structure. Such a representation can be utilized for applications such as indoor localization, where visual features solely do not adequately capture the scene’s structure. Though, graphs can encode the local structure of a scene by defining the relationships among the scene’s salient regions, yet, they are limited to only pairwise relationships. In contrast, hypergraphs can capture global structure, such as high-level relationships among regions effectively, and, thus, facilitate multimodal feature integration. Unlike existing multimodal scene representation with k-uniform hypergraphs, we represent multimodal scenes via a multimodal hypergraph that encodes the scene features at different levels and captures the scene’s global structure. The visual, geometric, and textual modalities of the panoramic 3D scene define the hyperedges (inter- and intra-relationship) among hypernodes (i.e., salient objects and texts). Additionally, we present an algorithm for sub-hypergraph inference that measures similarities between hypergraphs at global and local feature levels to facilitate the localization of panoramic images in indoor environments. The experimental results demonstrate the effectiveness of the proposed multimodal scene representation in indoor localization tasks.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multimodal Scene Representation using Hypergraph & Map Similarity Measure for Image Localization

  • Preeti Meena,
  • Himanshu Kumar,
  • Sandeep Yadav

摘要

Effective scene representation requires to encompass a complete understanding of the scene’s structure, including multimodal structure. Such a representation can be utilized for applications such as indoor localization, where visual features solely do not adequately capture the scene’s structure. Though, graphs can encode the local structure of a scene by defining the relationships among the scene’s salient regions, yet, they are limited to only pairwise relationships. In contrast, hypergraphs can capture global structure, such as high-level relationships among regions effectively, and, thus, facilitate multimodal feature integration. Unlike existing multimodal scene representation with k-uniform hypergraphs, we represent multimodal scenes via a multimodal hypergraph that encodes the scene features at different levels and captures the scene’s global structure. The visual, geometric, and textual modalities of the panoramic 3D scene define the hyperedges (inter- and intra-relationship) among hypernodes (i.e., salient objects and texts). Additionally, we present an algorithm for sub-hypergraph inference that measures similarities between hypergraphs at global and local feature levels to facilitate the localization of panoramic images in indoor environments. The experimental results demonstrate the effectiveness of the proposed multimodal scene representation in indoor localization tasks.