Classifier Recalibration for Human-Object Interaction Detection
摘要
Human-Object Interaction (HOI) detection aims to localize human-object pairs and recognize their interactions. However, existing HOI datasets suffer from long-tail distribution issues, which make recognizing rare interaction categories challenging. In similar visual tasks, supervised contrastive learning has shown potential in suppressing long-tail bias by effectively utilizing negative samples. Inspired by this, we propose a novel contrastive learning approach to recalibrate the biased HOI classifier in a fully trained model. We first introduce a hard negative mining technique to filter high-quality, easily misclassified HOI instance features from existing models. Then we propose a supervised prototypical contrastive learning method to recalibrate the HOI classifier using the hard negative features. Our method is efficient, widely applicable, and compatible with existing long-tail mitigation approaches, achieving significant improvements on three representative HOI baselines and two widely used HOI benchmarks (HICO-DET and V-COCO).