MAHC: motion-appearance video object segmentation via hierarchical attention and multi-level clustering
摘要
Unsupervised video object segmentation (UVOS) aims to automatically segment the most salient and semantically meaningful objects in a video without relying on manual annotations. Existing methods often focus on direct feature fusion without fully exploiting the inherent advantages of individual features, leading to limited performance when handling scenes with diverse motion patterns or similar foreground-background appearances. To address these challenges, we propose MAHC (motion-appearance hierarchical clustering), a novel unsupervised video object segmentation framework that effectively integrates motion and appearance cues through hierarchical feature learning and progressive clustering refinement. Our framework employs a hierarchical interleaved attention mechanism within an autoencoder structure to enhance motion feature representation, and utilizes a multi-level clustering strategy that progressively integrates different clustering techniques to achieve comprehensive segmentation from global patterns to fine-grained details. Additionally, we improve foreground-background discrimination by combining multi-view subspace analysis with motion intensity information from reconstructed optical flow. Extensive experiments on four challenging benchmark datasets (DAVIS16, FBMS, SegTrackv2, and YouTube-Objects) demonstrate that MAHC significantly outperforms existing unsupervised methods.