<p>Thanks to capability to alleviate the cost of large-scale annotation, few-shot action recognition (FSAR) has attracted increased attention of researchers in recent years. Existing FSAR approaches typically neglect the role of individual motion pattern in comparison, and under-explore the feature statistics for video dynamics. Thereby, they struggle to handle the challenging temporal misalignment in video dynamics, particularly by using 2D backbones. To overcome these limitations, this work proposes an adaptively aligned multi-scale second-order moment network, namely <InlineEquation ID="IEq3"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11263_2025_2432_Article_IEq1.gif" Format="GIF" Height="17" Rendition="HTML" Resolution="72" Type="Linedraw" Width="41" /> </InlineMediaObject> <EquationSource Format="TEX">\(\text {A}^2\text {M}^2\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <msup> <mtext>A</mtext> <mn>2</mn> </msup> <msup> <mtext>M</mtext> <mn>2</mn> </msup> </mrow> </math></EquationSource> </InlineEquation>-Net, to describe the latent video dynamics with a collection of powerful representation candidates and adaptively align them in an instance-guided manner. To this end, our <InlineEquation ID="IEq4"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11263_2025_2432_Article_IEq1.gif" Format="GIF" Height="17" Rendition="HTML" Resolution="72" Type="Linedraw" Width="41" /> </InlineMediaObject> <EquationSource Format="TEX">\(\text {A}^2\text {M}^2\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <msup> <mtext>A</mtext> <mn>2</mn> </msup> <msup> <mtext>M</mtext> <mn>2</mn> </msup> </mrow> </math></EquationSource> </InlineEquation>-Net involves two core components, namely, adaptive alignment (<InlineEquation ID="IEq5"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11263_2025_2432_Article_IEq5.gif" Format="GIF" Height="17" Rendition="HTML" Resolution="72" Type="Linedraw" Width="20" /> </InlineMediaObject> <EquationSource Format="TEX">\(\text {A}^2\)</EquationSource> <EquationSource Format="MATHML"><math> <msup> <mtext>A</mtext> <mn>2</mn> </msup> </math></EquationSource> </InlineEquation> module) for matching, and multi-scale second-order moment (<InlineEquation ID="IEq6"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11263_2025_2432_Article_IEq6.gif" Format="GIF" Height="17" Rendition="HTML" Resolution="72" Type="Linedraw" Width="22" /> </InlineMediaObject> <EquationSource Format="TEX">\(\text {M}^2\)</EquationSource> <EquationSource Format="MATHML"><math> <msup> <mtext>M</mtext> <mn>2</mn> </msup> </math></EquationSource> </InlineEquation> block) for strong representation. Specifically, <InlineEquation ID="IEq7"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11263_2025_2432_Article_IEq6.gif" Format="GIF" Height="17" Rendition="HTML" Resolution="72" Type="Linedraw" Width="22" /> </InlineMediaObject> <EquationSource Format="TEX">\(\text {M}^2\)</EquationSource> <EquationSource Format="MATHML"><math> <msup> <mtext>M</mtext> <mn>2</mn> </msup> </math></EquationSource> </InlineEquation> block develops a collection of semantic second-order descriptors at multiple spatio-temporal scales. Furthermore, <InlineEquation ID="IEq8"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11263_2025_2432_Article_IEq5.gif" Format="GIF" Height="17" Rendition="HTML" Resolution="72" Type="Linedraw" Width="20" /> </InlineMediaObject> <EquationSource Format="TEX">\(\text {A}^2\)</EquationSource> <EquationSource Format="MATHML"><math> <msup> <mtext>A</mtext> <mn>2</mn> </msup> </math></EquationSource> </InlineEquation> module aims to adaptively select informative candidate descriptors while considering the individual motion pattern. By such means, our <InlineEquation ID="IEq9"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11263_2025_2432_Article_IEq1.gif" Format="GIF" Height="17" Rendition="HTML" Resolution="72" Type="Linedraw" Width="41" /> </InlineMediaObject> <EquationSource Format="TEX">\(\text {A}^2\text {M}^2\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <msup> <mtext>A</mtext> <mn>2</mn> </msup> <msup> <mtext>M</mtext> <mn>2</mn> </msup> </mrow> </math></EquationSource> </InlineEquation>-Net is able to handle the challenging temporal misalignment problem by establishing an adaptive alignment protocol for strong representation. Notably, our proposed method generalizes well to various few-shot settings and diverse metrics. The experiments are conducted on five widely used FSAR benchmarks, and the results show our <InlineEquation ID="IEq10"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11263_2025_2432_Article_IEq1.gif" Format="GIF" Height="17" Rendition="HTML" Resolution="72" Type="Linedraw" Width="41" /> </InlineMediaObject> <EquationSource Format="TEX">\(\text {A}^2\text {M}^2\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <msup> <mtext>A</mtext> <mn>2</mn> </msup> <msup> <mtext>M</mtext> <mn>2</mn> </msup> </mrow> </math></EquationSource> </InlineEquation>-Net achieves very competitive performance compared to state-of-the-arts, demonstrating its effectiveness and generalization.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

\(\text {A}^2\text {M}^2\)-Net: Adaptively Aligned Multi-scale Moment for Few-Shot Action Recognition

  • Zilin Gao,
  • Qilong Wang,
  • Bingbing Zhang,
  • Qinghua Hu,
  • Peihua Li

摘要

Thanks to capability to alleviate the cost of large-scale annotation, few-shot action recognition (FSAR) has attracted increased attention of researchers in recent years. Existing FSAR approaches typically neglect the role of individual motion pattern in comparison, and under-explore the feature statistics for video dynamics. Thereby, they struggle to handle the challenging temporal misalignment in video dynamics, particularly by using 2D backbones. To overcome these limitations, this work proposes an adaptively aligned multi-scale second-order moment network, namely \(\text {A}^2\text {M}^2\) A 2 M 2 -Net, to describe the latent video dynamics with a collection of powerful representation candidates and adaptively align them in an instance-guided manner. To this end, our \(\text {A}^2\text {M}^2\) A 2 M 2 -Net involves two core components, namely, adaptive alignment ( \(\text {A}^2\) A 2 module) for matching, and multi-scale second-order moment ( \(\text {M}^2\) M 2 block) for strong representation. Specifically, \(\text {M}^2\) M 2 block develops a collection of semantic second-order descriptors at multiple spatio-temporal scales. Furthermore, \(\text {A}^2\) A 2 module aims to adaptively select informative candidate descriptors while considering the individual motion pattern. By such means, our \(\text {A}^2\text {M}^2\) A 2 M 2 -Net is able to handle the challenging temporal misalignment problem by establishing an adaptive alignment protocol for strong representation. Notably, our proposed method generalizes well to various few-shot settings and diverse metrics. The experiments are conducted on five widely used FSAR benchmarks, and the results show our \(\text {A}^2\text {M}^2\) A 2 M 2 -Net achieves very competitive performance compared to state-of-the-arts, demonstrating its effectiveness and generalization.