Illegal waste dumping, like all forms of pollution, can have detrimental effects on the environment and human well-being. The majority of current studies concentrate on garbage classification. While the recognition of littering behavior is rarely found in the existing works. In this study, we present a novel approach to identify illegal disposal of waste in surveillance footage recorded by on-site cameras. The illegal littering is recognized by a combined model. Convolutional processes in convolutional neural networks perform well in extracting local information but struggle to capture global representations. Multi-head self-attention for Vision Transformer is capable of capturing feature dependencies at large distances, but it can also destroy local feature details. This leads us to propose a new model based on MobileNet-v2 and Vision Transformer that can combine the benefits of Vision Transformer and CNNs. Furthermore, we use temporal channel attention in our method to improve the model’s capacity for temporal information interaction. We conducted experiments on our own dataset to validate the effectiveness of the proposed approach. The experimental results demonstrate that our model achieves good performance with the validation accuracy of 92.71%, obtaining the performance close to or even comparable with the state-of-the-art methods.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Combined Model Based on Deep Learning for Littering Detection

  • Cu Vinh Loc,
  • Truong Xuan Viet,
  • Tran Hoang Viet,
  • Le Hoang Thao,
  • Nguyen Hoang Viet

摘要

Illegal waste dumping, like all forms of pollution, can have detrimental effects on the environment and human well-being. The majority of current studies concentrate on garbage classification. While the recognition of littering behavior is rarely found in the existing works. In this study, we present a novel approach to identify illegal disposal of waste in surveillance footage recorded by on-site cameras. The illegal littering is recognized by a combined model. Convolutional processes in convolutional neural networks perform well in extracting local information but struggle to capture global representations. Multi-head self-attention for Vision Transformer is capable of capturing feature dependencies at large distances, but it can also destroy local feature details. This leads us to propose a new model based on MobileNet-v2 and Vision Transformer that can combine the benefits of Vision Transformer and CNNs. Furthermore, we use temporal channel attention in our method to improve the model’s capacity for temporal information interaction. We conducted experiments on our own dataset to validate the effectiveness of the proposed approach. The experimental results demonstrate that our model achieves good performance with the validation accuracy of 92.71%, obtaining the performance close to or even comparable with the state-of-the-art methods.