Collaborative learning of multimodal fusion and audio-guided attention for fish feeding behavior analysis in aquafarm
摘要
Efficient and accurate fish feeding is critical for enhancing aquatic productivity, reducing labor demands, and minimizing operational costs in modern aquaculture, thereby advancing the development of intelligent aquaculture systems. However, conventional feeding methods—reliant on manual expertise or rigid mechanical dispensers—fail to achieve adaptive and precision-based feeding control. To address this limitation, a novel Cross-Modal Hierarchical Gating (HGCM) network for fish feeding intensity recognition is proposed. The HGCM framework integrates an audio-guided attention module to dynamically prioritize salient image regions, while a hybrid architecture combining bidirectional LSTM and transformer encoder captures spatiotemporal feature dependencies. Further, through residual connections, gating mechanisms, and hierarchical multi-head attention, the model achieves deep cross-modal fusion of visual and auditory data. Experimental validation in commercial-scale aquaculture environments demonstrates the HGCM model’s superior performance, attaining 97.56% and 97.61% accuracy in triple-class feeding intensity classification on the Pseudocaranx dentex dataset and turbot dataset, respectively—significantly outperforming traditional unimodal approaches. The proposed system exhibits robust accuracy, noise resilience, and stability under challenging recirculating water conditions, offering a practical and scalable solution for intelligent feeding management. This work bridges a critical gap between theoretical research and industrial application, paving the way for data-driven, sustainable aquaculture practices.