ULLN: unifying local–long range context network for vehicle re-identification
摘要
The high similarity problem between vehicle appearances is a significant challenge that restricts the development of vehicle re-identification (Re-ID) tasks. The self-attention mechanism uses the dependency between pairs of pixels to capture the global context information of an image. However, some studies have shown that capturing fine-grained information from localized regions is more effective in solving the vehicle high similarity problem. In this paper, we explore how to use local context and global context modeling to more efficiently learn part-level features of vehicles and the overall appearance features of vehicles. We propose a unifying local-long range context network (ULLN) for vehicle Re-ID. The network performs self-attention computations at different scales in multiple dimensions to learn fine-grained information about localized areas of vehicles and structural information about vehicles overall. In ULLN, a unifying spatial local-long range context module (ULL-S) and a unifying channel local-long range context module (ULL-C) are designed to model the local context and the long range context from the spatial dimension and the channel dimension, respectively. Moreover, spatial (channel) inverse ratio constraint are introduced to realize implicit communication within the inter-window (interval) and the inter-grid (grill) regions. We further realize explicit communication within the inter-window (interval) and inter-grid (grill) by fusing the local context and the long range context. This fusion operation ensures that potential complementary and synergistic effects between local and global information are not overlooked. The results of experiments with three major public datasets, VeRi-776, VehicleID and VERI-Wild, show that our proposed ULLN achieves state-of-the-art performance.