<p>Zero-shot learning aims to transfer knowledge from seen to unseen classes, enabling the recognition of categories not present during training. However, two major challenges persist: feature redundancy and sample homogeneity. Extracted visual features contain semantically irrelevant information, which can hinder classification performance. In addition, due to the limited number of samples available for learning, the generated unseen class samples tend to be generic, lacking unique discriminative styles and diversity. To address these issues, we propose a novel unified framework called Generalized Zero-shot Learning based on Style and Feature Reconstruction (SFR-GZSL). The framework features a clear separation between feature disentanglement (extracting semantic content) and style reconstruction (handling appearance variations), with dedicated modules for each. Unlike existing hybrid methods that merely concatenate separate components, our key innovation lies in a unified framework that jointly optimizes feature disentanglement and style transfer through a novel OT-Attn module, which combines optimal transport and multi-head attention for precise style alignment and diverse sample generation. This innovative design ensures that the generated samples are more diverse and avoids uniformity, making it highly suitable for zero-shot learning scenarios. Our approach achieves the best results on four datasets (CUB, AWA2, FLO, and aPY) and more competitive results on the dataset SUN.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Generalized zero-shot learning based on style and feature reconstruction

  • Weichao Kong,
  • Chao Tan,
  • Siwei Chen,
  • Genlin Ji

摘要

Zero-shot learning aims to transfer knowledge from seen to unseen classes, enabling the recognition of categories not present during training. However, two major challenges persist: feature redundancy and sample homogeneity. Extracted visual features contain semantically irrelevant information, which can hinder classification performance. In addition, due to the limited number of samples available for learning, the generated unseen class samples tend to be generic, lacking unique discriminative styles and diversity. To address these issues, we propose a novel unified framework called Generalized Zero-shot Learning based on Style and Feature Reconstruction (SFR-GZSL). The framework features a clear separation between feature disentanglement (extracting semantic content) and style reconstruction (handling appearance variations), with dedicated modules for each. Unlike existing hybrid methods that merely concatenate separate components, our key innovation lies in a unified framework that jointly optimizes feature disentanglement and style transfer through a novel OT-Attn module, which combines optimal transport and multi-head attention for precise style alignment and diverse sample generation. This innovative design ensures that the generated samples are more diverse and avoids uniformity, making it highly suitable for zero-shot learning scenarios. Our approach achieves the best results on four datasets (CUB, AWA2, FLO, and aPY) and more competitive results on the dataset SUN.