Occlusion-aware vehicle re-identification via sparse token modeling under multi-perspective cues
摘要
Vehicle re-identification (Re-ID) is pivotal for intelligent transportation systems but remains challenging under severe occlusions and viewpoint variations, where existing models often fail by either incorporating distracting background noise or losing critical identity (ID)-specific information. To address the limitations of current datasets, restricted viewpoints, limited scene diversity, and insufficient occlusion samples, this paper introduces two new benchmarks: Vehicle-Syn, a large-scale synthetic dataset generated via game engines to simulate diverse weather and illumination conditions, and Occluded-Wild, a real-world dataset constructed upon VeRI-Wild to specifically evaluate occlusion robustness. Motivated by the need for robust feature purification and complementation, we propose a novel Occlusion-Aware Re-ID framework (OARI) that operates on a progressive logic of filtering, retrieval, and fusion. Our framework employs a sparse feature encoder to dynamically discard non-essential image tokens related to occlusions, followed by a feature matching module that retrieves K neighboring gallery images using joint image-level and patch-level similarity. Finally, a Transformer-based fusion module integrates complementary information from these neighbors to compensate for occlusion-induced feature loss. Experimental results demonstrate the effectiveness of our approach, achieving 75.2% Rank-1 accuracy on Occluded-Wild and 80.1% on Vehicle-Syn, yielding competitive performance compared to state-of-the-art methods under severe occlusion.