Memory Aggregation Network for Video Camouflaged Object Detection of Activated Sludge Microorganisms
摘要
Accurate detection of microorganisms in activated sludge is vital in wastewater treatment and environmental protection. However, existing methods ignore the fact that some microorganisms in wastewater microscopic videos are relatively small in size, and the target features are not obvious. Tracking the movement of larger microorganisms often requires frequent adjustments to the microscope’s field of view, which easily leads to blurring of the previous and next frames. Furthermore, the microbial information lacks obvious short-term temporal features and needs to rely on global contextual information for identification. To address these limitations, a Multi-scale Long-Term and Short-Term Memory Aggregation Network (MLSMAnet) is proposed based on ZoomNeXt, which is specifically used to detect camouflaged targets from wastewater microscopic videos. A Long-Term Feature Memory Aggregation Module is designed to enhance temporal consistency and discriminative representation of camouflaged targets. A Feature-Enhanced Sparse Attention Module and a Short-Term Weighted Fusion Group Unit are designed to improve the representational capability of small targets, suppress irrelevant features, and strengthen short-term inter-frame feature expression. A new dataset, EWMVCOD, is constructed for detecting camouflaged microorganisms from microscopic wastewater videos. Experiments on two benchmark datasets and our proposed EWMVCOD dataset demonstrate that the proposed MLSMAnet outperforms state-of-the-art video camouflaged object detection methods on EWMVCOD and exhibits competitive performance on the public benchmarks MoCA-Mask and CAD.