VLPSR: Enhancing Zero-Shot Object ReID with Vision-Language Model
摘要
As domains vary dramatically in the open world and re-training proves costly, zero-shot re-identification (ReID) emerges as a major challenge. Past works attempt to overfit on a single domain to achieve state-of-the-art in-domain baselines while failing to care about cross-domain adaptation. Our method, VLPSR, short for vision-language prompt with self-regularization, focuses on zero-shot learning, does not involve any additional training on the target domain, and exhibits a favorable performance with an over 4% accuracy increase on domain transfer, 2% on generalization and consistent in-domain performance, compared with CLIP-ReID, while saving numerous training hours.