Scene Graph Generation (SGG) forecasts relationships among entities in an image, vital for visual comprehension and reasoning tasks. Current SGG methods face three main challenges: limitations in multimodal integration, bias toward high-frequency predicate prediction, and training loss favoring seen triplets. These challenges lead to suboptimal performance in predicting zero-shot (unseen) triplets, limiting practical applications. To overcome these, we propose the Zero-Shot Scene Graph Generation with Bias Correction and Unseen Space Optimization (ZS-BUS) method. Specifically, to address multimodal fusion limitations, we introduce a Hybrid Attention Cross Network (HACNet) as an encoder to enhance intramodality refinement and intermodality interaction. To tackle high-frequency predicate bias, we utilize the Unseen Space Optimization Learning (USOL) module to eliminate irrelevant unseen triplets, focusing on plausible combinations. Lastly, to counter loss function bias toward seen triplets, we introduce a Progressive Refinement Correction Loss (PRCL), which regularizes the representation of infrequent triplets and enables the recognition of novel triplet combinations within partially annotated training scene graphs. Experimental evaluations confirm that ZS-BUS consistently outperforms current leading methods in the context of zero-shot SGG.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Zero-Shot Scene Graph Generation with Bias Correction and Unseen Space Optimization

  • Xinyue Li,
  • Yinsai Guo,
  • Liyan Ma,
  • Shaorong Xie

摘要

Scene Graph Generation (SGG) forecasts relationships among entities in an image, vital for visual comprehension and reasoning tasks. Current SGG methods face three main challenges: limitations in multimodal integration, bias toward high-frequency predicate prediction, and training loss favoring seen triplets. These challenges lead to suboptimal performance in predicting zero-shot (unseen) triplets, limiting practical applications. To overcome these, we propose the Zero-Shot Scene Graph Generation with Bias Correction and Unseen Space Optimization (ZS-BUS) method. Specifically, to address multimodal fusion limitations, we introduce a Hybrid Attention Cross Network (HACNet) as an encoder to enhance intramodality refinement and intermodality interaction. To tackle high-frequency predicate bias, we utilize the Unseen Space Optimization Learning (USOL) module to eliminate irrelevant unseen triplets, focusing on plausible combinations. Lastly, to counter loss function bias toward seen triplets, we introduce a Progressive Refinement Correction Loss (PRCL), which regularizes the representation of infrequent triplets and enables the recognition of novel triplet combinations within partially annotated training scene graphs. Experimental evaluations confirm that ZS-BUS consistently outperforms current leading methods in the context of zero-shot SGG.