Attribute-Based Out-of-Distribution Detection Using LLaVA
摘要
Deep neural networks (DNNs) often exhibit overconfidence when encountering out-of-distribution (OOD) samples, which poses significant challenges in real-world applications. To address this issue, large multimodal models (LMMs) have been employed, showing considerable promise. Existing approaches attempt to explore CLIP’s textual capabilities by generating extensive (OOD) categories. Recognizing that distinctive attributes of various image categories are essential for differentiating between in-distribution (ID) and OOD samples, this paper introduces an attribute-based method for OOD detection. This approach utilizes the LLaVA to extract image attributes, which are then compared with a reference attribute set established for each ID category to estimate the likelihood of an image being ID or OOD. Furthermore, to comprehensively represent each category, we introduce an attribute selection strategy that considers both the commonality and diversity of attributes, significantly improving OOD detection performance. Enhancing OOD detection performance. Extensive experiments conducted across various ID/OOD settings demonstrate the effectiveness of our method and its superiority over state-of-the-art approaches.