Open-World Object Detection (OWOD) is designed to detect known and unknown objects in evolving scenarios. A persistent issue in this task is the incorrect classification of unfamiliar objects as background. While background areas usually conform to a straightforward, unimodal statistical pattern, object instances in the foreground are often diverse and follow long-tailed distributions. Existing approaches frequently fail to consider this discrepancy by employing a one-size-fits-all modeling strategy. To tackle this limitation, we propose Decoupled Modeling of Foreground and Background (DMFB). At the heart of this framework is Reconstruction Error-based Decoupled Modeling (REDM), which separates foreground and background through reconstruction error analysis and uses multi-scale strategies to better model their respective traits. Furthermore, we propose the Unsupervised Proposal Generation Module (UPM), which utilizes knowledge transfer from a large visual model (LVM) to generate pseudo-labels, while a dual-filtering process helps reduce conflicts between noisy and true labels. Comprehensive experiments on OWOD benchmark show that DMFB improves unknown object detection by 40.3% (reaching 53.2 U-Recall), surpassing state-of-the-art methods.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Decoupled Modeling of Foreground and Background for Open-World Object Detection

  • Linhua Ye,
  • Xing Xi,
  • Yangyang Huang,
  • Ronghua Luo

摘要

Open-World Object Detection (OWOD) is designed to detect known and unknown objects in evolving scenarios. A persistent issue in this task is the incorrect classification of unfamiliar objects as background. While background areas usually conform to a straightforward, unimodal statistical pattern, object instances in the foreground are often diverse and follow long-tailed distributions. Existing approaches frequently fail to consider this discrepancy by employing a one-size-fits-all modeling strategy. To tackle this limitation, we propose Decoupled Modeling of Foreground and Background (DMFB). At the heart of this framework is Reconstruction Error-based Decoupled Modeling (REDM), which separates foreground and background through reconstruction error analysis and uses multi-scale strategies to better model their respective traits. Furthermore, we propose the Unsupervised Proposal Generation Module (UPM), which utilizes knowledge transfer from a large visual model (LVM) to generate pseudo-labels, while a dual-filtering process helps reduce conflicts between noisy and true labels. Comprehensive experiments on OWOD benchmark show that DMFB improves unknown object detection by 40.3% (reaching 53.2 U-Recall), surpassing state-of-the-art methods.