The integration of machine learning (ML) to predict the Biotic Index (BI) for assessing Ecological Quality (EQ) of marine environments using environmental DNA (eDNA) metabarcoding data represents a significant advance in biomonitoring. However, the complexity of these data types poses challenges like systematic variability and the curse of dimensionality, impacting prediction quality. To address these issues, we propose a generic ML pipeline with crucial preprocessing steps, including normalization and dimensionality reduction techniques, to enhance prediction quality. Our work involves comparing the performance of the Random Forest Classifier for predicting EQ classes across seven markers. We examine various combinations of two normalization techniques and a suitable dimensionality reduction with an optimal number of reduced features. Through rigorous experimentation, we identify the most effective combination and optimal number of components to retain, establishing a standardized preprocessing protocol for this data type. This robust framework for preprocessing metabarcoding data significantly advances EQ prediction.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Efficient Data Preprocessing for Ecological Quality Assessment in Marine Environments

  • Houria Braikia,
  • Sana Ben Hamida,
  • Marta Rukoz

摘要

The integration of machine learning (ML) to predict the Biotic Index (BI) for assessing Ecological Quality (EQ) of marine environments using environmental DNA (eDNA) metabarcoding data represents a significant advance in biomonitoring. However, the complexity of these data types poses challenges like systematic variability and the curse of dimensionality, impacting prediction quality. To address these issues, we propose a generic ML pipeline with crucial preprocessing steps, including normalization and dimensionality reduction techniques, to enhance prediction quality. Our work involves comparing the performance of the Random Forest Classifier for predicting EQ classes across seven markers. We examine various combinations of two normalization techniques and a suitable dimensionality reduction with an optimal number of reduced features. Through rigorous experimentation, we identify the most effective combination and optimal number of components to retain, establishing a standardized preprocessing protocol for this data type. This robust framework for preprocessing metabarcoding data significantly advances EQ prediction.