<p>Audio-visual zero-shot learning requires an understanding of the relationship between audio and visual information to determine unseen classes. Despite many efforts and significant progress in the field, many existing methods tend to focus on learning strong representations, neglecting the semantic consistency between audio and video as well as the inherent hierarchical structure of the data. To address these issues, we propose Learning Semantic Consistency for Audio-Visual Zero-shot Learning. Specifically, we employ an attention mechanism to enhance cross-modal information interactions, aiming to capture the semantic consistency between audio and visual data. Meanwhile, we introduce a hyperbolic space to model the hierarchical structure of the data itself. Moreover, the proposed approach includes a novel loss function that considers the relationships between input modalities, reducing the distance between features of different modalities. To evaluate the proposed method, we test it on three benchmark datasets <InlineEquation ID="IEq1"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="10462_2025_11228_Article_IEq1.gif" Format="GIF" Height="18" Rendition="HTML" Resolution="72" Type="Linedraw" Width="145" /> </InlineMediaObject> <EquationSource Format="TEX">\(\hbox {VGGSound-GZS}{{\textrm{L}}^{cls}}\)</EquationSource> </InlineEquation>, <InlineEquation ID="IEq2"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="10462_2025_11228_Article_IEq2.gif" Format="GIF" Height="18" Rendition="HTML" Resolution="72" Type="Linedraw" Width="99" /> </InlineMediaObject> <EquationSource Format="TEX">\(\hbox {UCF-GZS}{{\textrm{L}}^{cls}}\)</EquationSource> </InlineEquation>, and <InlineEquation ID="IEq3"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="10462_2025_11228_Article_IEq3.gif" Format="GIF" Height="21" Rendition="HTML" Resolution="72" Type="Linedraw" Width="147" /> </InlineMediaObject> <EquationSource Format="TEX">\(\hbox {ActivityNet-GZS}{{\textrm{L}}^{cls}}\)</EquationSource> </InlineEquation>. Extensive experimental results show that the proposed method achieves state-of-the-art performance on all three datasets. For example, on the <InlineEquation ID="IEq4"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="10462_2025_11228_Article_IEq2.gif" Format="GIF" Height="18" Rendition="HTML" Resolution="72" Type="Linedraw" Width="99" /> </InlineMediaObject> <EquationSource Format="TEX">\(\hbox {UCF-GZS}{{\textrm{L}}^{cls}}\)</EquationSource> </InlineEquation> dataset, the harmonic mean is improved by 5.7%. Code and data available at <a href="https://github.com/ybyangjing/LSC-AVZSL">https://github.com/ybyangjing/LSC-AVZSL</a>.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Learning semantic consistency for audio-visual zero-shot learning

  • Xiaoyong Li,
  • Jing Yang,
  • Yuling Chen,
  • Wei Zhang,
  • Xiaoli Ruan,
  • Chengjiang Li,
  • Zhidong Su

摘要

Audio-visual zero-shot learning requires an understanding of the relationship between audio and visual information to determine unseen classes. Despite many efforts and significant progress in the field, many existing methods tend to focus on learning strong representations, neglecting the semantic consistency between audio and video as well as the inherent hierarchical structure of the data. To address these issues, we propose Learning Semantic Consistency for Audio-Visual Zero-shot Learning. Specifically, we employ an attention mechanism to enhance cross-modal information interactions, aiming to capture the semantic consistency between audio and visual data. Meanwhile, we introduce a hyperbolic space to model the hierarchical structure of the data itself. Moreover, the proposed approach includes a novel loss function that considers the relationships between input modalities, reducing the distance between features of different modalities. To evaluate the proposed method, we test it on three benchmark datasets \(\hbox {VGGSound-GZS}{{\textrm{L}}^{cls}}\) , \(\hbox {UCF-GZS}{{\textrm{L}}^{cls}}\) , and \(\hbox {ActivityNet-GZS}{{\textrm{L}}^{cls}}\) . Extensive experimental results show that the proposed method achieves state-of-the-art performance on all three datasets. For example, on the \(\hbox {UCF-GZS}{{\textrm{L}}^{cls}}\) dataset, the harmonic mean is improved by 5.7%. Code and data available at https://github.com/ybyangjing/LSC-AVZSL.