Interpretation and visualization of the behavior of detection transformers often highlight image areas attended to by the model but provide limited insight into the semantics that the model is focusing on. This paper introduces an extension to detection transformers that learns and utilizes prototypical local features for object detection, termed prototypical parts. These features are designed to be mutually exclusive and align with the detection classes of the model. The proposed extension consists of a bottleneck module, the prototype neck that computes a sparse representation in terms of prototype activations, and a novel loss term that aligns prototypes with object classes. This approach leads to interpretable representations in the prototype neck, allowing visual inspection of the image content as perceived by the model and a better understanding of model reliability. Results show that our method incurs only a limited performance penalty. Furthermore, we provide examples that showcase the explanatory power of our approach, justifying the slight performance reduction.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

ProtoP-OD: Explainable Object Detection with Prototypical Parts

  • Pavlos Rath-Manakidis,
  • Frederik Strothmann,
  • Tobias Glasmachers,
  • Laurenz Wiskott

摘要

Interpretation and visualization of the behavior of detection transformers often highlight image areas attended to by the model but provide limited insight into the semantics that the model is focusing on. This paper introduces an extension to detection transformers that learns and utilizes prototypical local features for object detection, termed prototypical parts. These features are designed to be mutually exclusive and align with the detection classes of the model. The proposed extension consists of a bottleneck module, the prototype neck that computes a sparse representation in terms of prototype activations, and a novel loss term that aligns prototypes with object classes. This approach leads to interpretable representations in the prototype neck, allowing visual inspection of the image content as perceived by the model and a better understanding of model reliability. Results show that our method incurs only a limited performance penalty. Furthermore, we provide examples that showcase the explanatory power of our approach, justifying the slight performance reduction.