Long-tailed video recognition via majority-guided diffusion model
摘要
Long-tailed video recognition presents a significant challenge due to the imbalanced distribution of samples across different classes, where majority classes contain abundant samples, and minority classes are severely underrepresented. Existing methods primarily focus on resampling, reweighting, and architectural modifications, which do not fully leverage the rich information contained within majority class samples. Motivated by this, we propose a novel majority-guided diffusion model for addressing the long-tailed distribution problem in video recognition. Specifically, we introduce an attention-based feature mix module (AFM) to blend majority and minority class information, followed by a minority-class data generator (MDG) that synthesizes diverse minority class samples using a latent diffusion model. By leveraging the rich information from majority class samples, our method generates realistic minority class samples that improve the overall model performance on underrepresented categories. Extensive experimental results on long-tailed video recognition benchmarks validate the effectiveness of the proposed framework.