Purpose <p>Hyperspectral sensing (remote or proximal) has emerged as a pivotal tool to classify plant materials (seeds, leaves, and whole plants), pharmaceutical products, food items, and many other objects. Thus, hyperspectral sensing is one of the most frequently used technologies in research articles published by this journal, and it was therefore found relevant to address two methodological issues, which (based on Google Scholar searches) appear to be over-looked or ignored in &gt;94% of hyperspectral sensing studies: 1) the "small N, large P" problem, when number of spectral bands (explanatory variables, “P”) surpasses number of observations, (“N”) leading to potential model over-fitting, and 2) absence of independent validation data in performance assessments of classification models.</p> Methods <p>Based on simulations of randomly generated data, risks associated with these issues were illustrated. This communication explores and discusses consequences of over-fitting and risks of misleadingly high accuracy that can result from having a large number of variables relative to observations. Moreover, connections of these issues with radiometric repeatability (levels of stochastic noise) are highlighted. A method is proposed wherein a theoretical dataset is generated to mirror the structure of an actual dataset, with the classification of this theoretical dataset serving as a reference.</p> Conclusion <p>By shedding light on important and common experimental design issues, the principal aim is to enhance methodological rigor and transparency in classifications of hyperspectral sensing data and foster improved and effective applications across various science domains, including precision agriculture.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Experimental design issues associated with classifications of hyperspectral sensing data

  • Christian Nansen,
  • Hyoseok Lee,
  • Mohsen B. Mesgaran

摘要

Purpose

Hyperspectral sensing (remote or proximal) has emerged as a pivotal tool to classify plant materials (seeds, leaves, and whole plants), pharmaceutical products, food items, and many other objects. Thus, hyperspectral sensing is one of the most frequently used technologies in research articles published by this journal, and it was therefore found relevant to address two methodological issues, which (based on Google Scholar searches) appear to be over-looked or ignored in >94% of hyperspectral sensing studies: 1) the "small N, large P" problem, when number of spectral bands (explanatory variables, “P”) surpasses number of observations, (“N”) leading to potential model over-fitting, and 2) absence of independent validation data in performance assessments of classification models.

Methods

Based on simulations of randomly generated data, risks associated with these issues were illustrated. This communication explores and discusses consequences of over-fitting and risks of misleadingly high accuracy that can result from having a large number of variables relative to observations. Moreover, connections of these issues with radiometric repeatability (levels of stochastic noise) are highlighted. A method is proposed wherein a theoretical dataset is generated to mirror the structure of an actual dataset, with the classification of this theoretical dataset serving as a reference.

Conclusion

By shedding light on important and common experimental design issues, the principal aim is to enhance methodological rigor and transparency in classifications of hyperspectral sensing data and foster improved and effective applications across various science domains, including precision agriculture.