<p>Although semantic segmentation shows promising results in Document Layout Analysis (DLA) tasks, it typically demands a large set of annotated images for training and faces difficulties in generalizing to unseen document categories. Few-shot learning addresses the challenges of limited document layout data, however, entanglement of semantic features between support and query images often impairs generalization and diminishes DLA performance. In this paper, we introduce the Few-Shot Quaternion-valued Correlation Squeeze Network (FS-QCSNet) to address these challenges using a limited number of annotated document images. Our method leverages quaternion convolution in correlation learning to jointly model interactions between support and query subspaces through Hamilton products, which inherently capture latent dependencies among multi-dimensional features. Compared to real-valued counterparts, quaternion convolution significantly reduces model complexity by sharing weights across feature dimensions, while preserving rich representation capabilities. To further refine discriminative layout patterns, we incorporate a parameter-free Simam attention module that prioritizes task-relevant regions guided by quaternion features. Finally, the Residual 2D Decoder reconstructs segmentation masks by effectively utilizing compact and interaction-aware features generated from quaternion operations. Remarkably, the proposed FS-QCSNet model achieves superior performance on real-world document image datasets, such as the DSSE-200 dataset, the Layout Analysis dataset, and the CDSSE dataset, surpassing previous few-shot semantic segmentation methods.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Few-Shot Quaternion-valued Correlation Squeeze Network for Document Image Layout Segmentation

  • Rui Yao,
  • Qiwei Yu,
  • Songhui Zhao,
  • Yuxuan Yang,
  • Yong Zhou,
  • Bing Liu

摘要

Although semantic segmentation shows promising results in Document Layout Analysis (DLA) tasks, it typically demands a large set of annotated images for training and faces difficulties in generalizing to unseen document categories. Few-shot learning addresses the challenges of limited document layout data, however, entanglement of semantic features between support and query images often impairs generalization and diminishes DLA performance. In this paper, we introduce the Few-Shot Quaternion-valued Correlation Squeeze Network (FS-QCSNet) to address these challenges using a limited number of annotated document images. Our method leverages quaternion convolution in correlation learning to jointly model interactions between support and query subspaces through Hamilton products, which inherently capture latent dependencies among multi-dimensional features. Compared to real-valued counterparts, quaternion convolution significantly reduces model complexity by sharing weights across feature dimensions, while preserving rich representation capabilities. To further refine discriminative layout patterns, we incorporate a parameter-free Simam attention module that prioritizes task-relevant regions guided by quaternion features. Finally, the Residual 2D Decoder reconstructs segmentation masks by effectively utilizing compact and interaction-aware features generated from quaternion operations. Remarkably, the proposed FS-QCSNet model achieves superior performance on real-world document image datasets, such as the DSSE-200 dataset, the Layout Analysis dataset, and the CDSSE dataset, surpassing previous few-shot semantic segmentation methods.