Quantifying the Complexity of Agricultural Data for Regression and Classification Problems
摘要
Agriculture is a crucial sector that sustains the global population, yet it faces increasing challenges due to climate change, resource limitations, and the need for sustainable practices. Machine learning (ML) has emerged as a powerful tool in addressing these challenges, particularly in tasks like crop yield prediction and crop recommendation. However, the effectiveness of ML models in agriculture heavily depends on the complexity of the underlying datasets. Understanding this complexity is essential for selecting appropriate algorithms and optimizing model performance. This chapter aims to quantify the complexity of agricultural datasets used in regression and classification problems. By leveraging a set of complexity measures, we assess the complexity of the crop yield prediction problem (regression) and the crop recommendation problem (classification). Our findings reveal significant variability in the complexity of different crops. For instance, yield prediction for crops like banana and coconut is relatively straightforward due to high feature correlations, while crops like castor present more complex challenges due to low correlation and high non-linearity in the data. In the classification context, crops such as grapes and apple exhibit lower complexity, making them easier to classify, whereas other crops pose greater difficulties. Overall, this chapter provides a detailed evaluation of agricultural data complexity, offering insights that can guide the selection and development of machine learning models for various agricultural tasks.