Mamba-Driven Comprehensive Context Learning for Zero-Shot HOI Detection
摘要
Human-Object Interaction (HOI) detection aims to simultaneously localize human-object pairs and identify their interactions, while the zero-shot setting imposes greater demands on the generalization capabilities. Existing state-of-the-art zero-shot HOI detection methods typically build on the Vision-Language Models (VLMs) paradigm, leveraging the rich semantic knowledge in VLMs to enhance the transferability of the models. However, the insufficient excavation of semantic and spatial information, together with the inherent challenge of heterogeneous feature fusion, severely constrains their overall performance and scalability. To address these issues, we propose a novel Mamba-Driven Comprehensive Context Learning approach (MCCHOI) for Zero-Shot HOI Detection. MCCHOI establishes a Mamba-based semantic-spatial collaborative learning architecture, which first bridges the gap between semantic and spatial feature spaces through bidirectional interaction context modeling, and then integrates semantic and spatial features under the guidance of a semantic-spatial synergy mechanism, constructing comprehensive and generalizable HOI representations. Specifically, to tackle feature extraction bottlenecks, we first introduce a spatial-aware scan mechanism that integrates broader regional context into the scanning process, enabling the adaptive aggregation of multi-scale spatial features to generate fine-grained spatial representations. Subsequently, we design a Mamba-based semantic-to-spatial refinement strategy equipped with a context-aware parameter generation mechanism, which excavates multi-view interaction cues under the guidance of global semantic-spatial knowledge to facilitate context propagation from the semantic domain to the spatial domain, thereby enhancing the discriminability of spatial features. In addition, we propose a Mamba-based spatial-to-semantic modulation strategy with dynamic prior augmentation, injecting external semantic-spatial priors into the visual semantics to facilitate context propagation from the spatial domain to the semantic domain, yielding interaction-aware semantic representations. Finally, we introduce a semantic-spatial synergy module, which explicitly models the complex interdependencies between semantic-spatial representations through intra-domain excavation and inter-domain fusion. This module constructs a robust and holistic interaction context, effectively addressing the challenges posed by heterogeneous feature integration in previous approaches. Extensive experiments on two benchmarks demonstrate the effectiveness of the proposed MCCHOI under the zero-shot setting.