<p>Language-based object detection aims to locate target objects from complex language queries. However, current vision-language detectors often struggle to understand complex representations of visual objects (<i>e</i>.<i>g</i>., attributes, shapes, and relationships), especially under complex queries. In this paper, we first conduct a thorough analysis of current language-based detectors to identify their specific weaknesses in compositional understanding. To this end, we propose a <i>novel comprehensive evaluation framework</i> that automatically categorizes test cases by the type and complexity of compositionality, leveraging large language models (LLMs). This reveals that detectors show significant performance drops with increased complexity and consistent failures in specific types, such as spatial and numerical reasoning. To effectively address this, we propose a multifaceted synthetic data consisting of (1) <i>generative model-based synthetic triplets</i> that inherited compositional knowledge from large generative models (e.g., LLMs, diffusion models) in the form of triplets (<i>i</i>.<i>e</i>., image-text-box data); and (2) <i>weakness-targeted synthetic descriptions</i> designed to enhance understanding in vulnerable types like spatial and numeracy concepts. We further introduce a <i>compositional contrastive learning</i> method to better leverage the proposed synthetic data while mitigating the common drawbacks of synthetic data. Consequently, our models trained on proposed multifaceted synthetic data exhibit a significant performance boost in the Omnilabel benchmark by up to +7.1AP and the <InlineEquation ID="IEq1"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11263_2025_2554_Article_IEq1.gif" Format="GIF" Height="17" Rendition="HTML" Resolution="72" Type="Linedraw" Width="20" /> </InlineMediaObject> <EquationSource Format="TEX">\(\hbox {D}^{3}\)</EquationSource> <EquationSource Format="MATHML"><math> <msup> <mtext>D</mtext> <mn>3</mn> </msup> </math></EquationSource> </InlineEquation> benchmark by up to <InlineEquation ID="IEq2"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11263_2025_2554_Article_IEq2.gif" Format="GIF" Height="14" Rendition="HTML" Resolution="72" Type="Linedraw" Width="35" /> </InlineMediaObject> <EquationSource Format="TEX">\(+8.4\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mo>+</mo> <mn>8.4</mn> </mrow> </math></EquationSource> </InlineEquation>AP upon existing baselines.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Learning Compositionality from Multifaceted Synthetic Data for Language-based Object Detection

  • Kwanyong Park,
  • Sojung An,
  • Yong Jae Lee,
  • Donghyun Kim

摘要

Language-based object detection aims to locate target objects from complex language queries. However, current vision-language detectors often struggle to understand complex representations of visual objects (e.g., attributes, shapes, and relationships), especially under complex queries. In this paper, we first conduct a thorough analysis of current language-based detectors to identify their specific weaknesses in compositional understanding. To this end, we propose a novel comprehensive evaluation framework that automatically categorizes test cases by the type and complexity of compositionality, leveraging large language models (LLMs). This reveals that detectors show significant performance drops with increased complexity and consistent failures in specific types, such as spatial and numerical reasoning. To effectively address this, we propose a multifaceted synthetic data consisting of (1) generative model-based synthetic triplets that inherited compositional knowledge from large generative models (e.g., LLMs, diffusion models) in the form of triplets (i.e., image-text-box data); and (2) weakness-targeted synthetic descriptions designed to enhance understanding in vulnerable types like spatial and numeracy concepts. We further introduce a compositional contrastive learning method to better leverage the proposed synthetic data while mitigating the common drawbacks of synthetic data. Consequently, our models trained on proposed multifaceted synthetic data exhibit a significant performance boost in the Omnilabel benchmark by up to +7.1AP and the \(\hbox {D}^{3}\) D 3 benchmark by up to \(+8.4\) + 8.4 AP upon existing baselines.