Saliency Based Data Augmentation for Few-Shot Video Action Recognition
摘要
Despite the progress made in few-shot video action recognition, existing methods still struggle to achieve satisfactory performance when support samples are limited (e.g., 1-shot task). This paper proposes to augment training samples without relying on additional supervision and labor costs, aiming at improving generalizability of learned representations. We introduce a novel self-supervised salient object detection model which results in frame-level saliency and background features of videos. A shared encoder is employed to fuse saliency and background information from different videos. Both intra- and inter-class fusion are performed, in which the latter is controlled by prior probability to avoid semantic ambiguities. This way actually corresponds to augment training data in feature space. The saliency-background representations formed from query and support videos are used to construct class prototypes through Temporal-Relational CrossTransformers. Experimental results on four standard benchmarks demonstrate that the proposed method outperforms state-of-the-arts under various few-shot settings, particularly excelling in the 1-shot case.