<p>Generating descriptions for fashion products is valuable as it helps customers understand their features and make informed purchasing decisions, thereby holding substantial economic significance. In reality, many fashion products are presented from multiple angles and have a complex hierarchical structure that encompasses various spatial granularities. Existing methods that capture product-invariant features are inadequate for extracting the discriminative information necessary to differentiate between products, as the information derived from invariant product features is limited in comparison to that obtained from the original features. Moreover, fashion datasets frequently exhibit a disparity between the information conveyed in images and captions. This discrepancy arises when datasets contain multiple angles all paired with a single description, resulting in situations where certain textual details may not be visible in any angle. To address such issues, we propose a novel <b>F</b>ashion <b>T</b>ransformer framework, which consists of two components: (1) an intra-group confusion module and an inter-group discrepancy module with dual feature interaction branches integrated into a unified model, and (2) a fashion encoder generates masks that identify the feature tokens and text tokens where information co-occurs in both modalities, ensuring precise alignment between the fashion images and texts. The experiments conducted on two standard benchmarks demonstrate that Fashion Transformer generates concept-centered descriptions that maintain high consistency with the ground-truth word. More remarkably, Fashion Transformer achieves competitive performance on FACAD and Fashion-Gen datasets, with the CIDEr-D score being increased from 67.2% to 70.8%, 190.7% to 201.6%, respectively.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Enhanced group relation learning via aligned attention masking for fashion product captioning

  • Yuhao Tang,
  • Dong Ye,
  • Fei Tao,
  • Guodong Du

摘要

Generating descriptions for fashion products is valuable as it helps customers understand their features and make informed purchasing decisions, thereby holding substantial economic significance. In reality, many fashion products are presented from multiple angles and have a complex hierarchical structure that encompasses various spatial granularities. Existing methods that capture product-invariant features are inadequate for extracting the discriminative information necessary to differentiate between products, as the information derived from invariant product features is limited in comparison to that obtained from the original features. Moreover, fashion datasets frequently exhibit a disparity between the information conveyed in images and captions. This discrepancy arises when datasets contain multiple angles all paired with a single description, resulting in situations where certain textual details may not be visible in any angle. To address such issues, we propose a novel Fashion Transformer framework, which consists of two components: (1) an intra-group confusion module and an inter-group discrepancy module with dual feature interaction branches integrated into a unified model, and (2) a fashion encoder generates masks that identify the feature tokens and text tokens where information co-occurs in both modalities, ensuring precise alignment between the fashion images and texts. The experiments conducted on two standard benchmarks demonstrate that Fashion Transformer generates concept-centered descriptions that maintain high consistency with the ground-truth word. More remarkably, Fashion Transformer achieves competitive performance on FACAD and Fashion-Gen datasets, with the CIDEr-D score being increased from 67.2% to 70.8%, 190.7% to 201.6%, respectively.