Compared to conventional RGB semantic segmentation methods, light field semantic segmentation approaches can incorporate more fine-grained scene information, often leading to superior performance. However, the high-dimensional characteristic of the light field also introduces numerous redundant information and computational burdens. In this paper, we propose a novel light field semantic segmentation network named Light Field Efficient Aggregation Network(LF-EANet), which can efficiently aggregate the structured scene information in light field and generate more precise scene semantic understanding. On the one hand, we encode the spatial position information of light field Sub-Aperture Images(SAIs) to ensure the incorporation of the most valuable perspective position information. On the other hand, we explore multi-scale feature interaction mechanisms from spatial and angular levels to generate more representative auxiliary scene feature. Ultimately, our network provides the center view image with robust scene feature augmentation and semantic perception guidance. This method shows excellent performance on both real-world and synthetic light field semantic segmentation dataset.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multi-scale Spatial-Angular Information Aggregation Network for Image Semantic Segmentation

  • Yiming Li,
  • Ruixuan Cong,
  • Sizhe Wang,
  • Jiahao Shen,
  • Hao Sheng

摘要

Compared to conventional RGB semantic segmentation methods, light field semantic segmentation approaches can incorporate more fine-grained scene information, often leading to superior performance. However, the high-dimensional characteristic of the light field also introduces numerous redundant information and computational burdens. In this paper, we propose a novel light field semantic segmentation network named Light Field Efficient Aggregation Network(LF-EANet), which can efficiently aggregate the structured scene information in light field and generate more precise scene semantic understanding. On the one hand, we encode the spatial position information of light field Sub-Aperture Images(SAIs) to ensure the incorporation of the most valuable perspective position information. On the other hand, we explore multi-scale feature interaction mechanisms from spatial and angular levels to generate more representative auxiliary scene feature. Ultimately, our network provides the center view image with robust scene feature augmentation and semantic perception guidance. This method shows excellent performance on both real-world and synthetic light field semantic segmentation dataset.