<p>Automatic classification models of pelagic species based on trawl–acoustic survey data link echotrace characteristics observed in echosounder echograms to species caught in pelagic trawl hauls. In multispecies ecosystems, developing such models is particularly challenging because pelagic trawls provide probabilistic, proportion‑based species information rather than unambiguous labels for individual schools, which limits the applicability of strictly supervised training and motivates weakly-supervised (probabilistic) learning approaches. This study aims to improve the performance of a previously developed weakly-supervised multi-output classification model through a conservative self-training technique. Using the model’s probabilistic multispecies outputs, self-training gradually transforms the most confident predictions among the poorly-labelled cases into pseudo-labelled ones, increasing the ratio of labelled data and strengthening both model learning and evaluation for the next iterations. The approach was tested on a multi-output classification model trained on small pelagic species in the Bay of Biscay. Self-training improved overall out-of-sample accuracy from 63.5% to 72% (F1-score from 53.7% to 60.7%), maintaining or improving performance for each individual species. Notably, it substantially enhanced performance for species that are abundant but rarely captured in monospecific trawls, such as sardine, whose accuracy increased from 50% to 86% (F1-score from 27.3% to 52.7%). In contrast, performance for species with enough labelled cases remained stable throughout the process.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A self-training approach to improve multispecies school classification models in fisheries acoustics data

  • Aitor Lekanda,
  • Guillermo Boyra,
  • Nils Olav Handegard,
  • Maite Louzao

摘要

Automatic classification models of pelagic species based on trawl–acoustic survey data link echotrace characteristics observed in echosounder echograms to species caught in pelagic trawl hauls. In multispecies ecosystems, developing such models is particularly challenging because pelagic trawls provide probabilistic, proportion‑based species information rather than unambiguous labels for individual schools, which limits the applicability of strictly supervised training and motivates weakly-supervised (probabilistic) learning approaches. This study aims to improve the performance of a previously developed weakly-supervised multi-output classification model through a conservative self-training technique. Using the model’s probabilistic multispecies outputs, self-training gradually transforms the most confident predictions among the poorly-labelled cases into pseudo-labelled ones, increasing the ratio of labelled data and strengthening both model learning and evaluation for the next iterations. The approach was tested on a multi-output classification model trained on small pelagic species in the Bay of Biscay. Self-training improved overall out-of-sample accuracy from 63.5% to 72% (F1-score from 53.7% to 60.7%), maintaining or improving performance for each individual species. Notably, it substantially enhanced performance for species that are abundant but rarely captured in monospecific trawls, such as sardine, whose accuracy increased from 50% to 86% (F1-score from 27.3% to 52.7%). In contrast, performance for species with enough labelled cases remained stable throughout the process.