This article focuses on clustering 514 cities/districts in Indonesia based on their poverty levels from the pre-COVID-19 period (2019) until the near-end of the COVID-19 period (2022). Using publicly available data from the Central Statistics Agency (BPS) for 2019–2022, encompassing 11 poverty-related variables, the study employs unsupervised learning techniques. The process involves testing for multicollinearity; normalizing the data and applying Principal Component Analysis to reduce the variables to 2 principal components; determining the optimal number of clusters (3) is done using the elbow method, silhouette index, and gap statistic algorithm; the K-Means algorithm is then applied to cluster the cities/districts each year. The analysis reveals three clusters: cluster 1 represents areas with low poverty levels, cluster 2 comprises regions with moderate levels, and cluster 3 includes areas with high poverty rates. The number of cities/districts in cluster 3 increased from 36 in 2019 to 40 in 2022, with 32 consistently belonging to this cluster. These findings highlight the importance of targeted government assistance to these areas.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Cluster Analysis of Poverty Data in Cities/Districts in Indonesia Using K-Means Algorithm for the Years 2019–2022

  • Julian Salomo,
  • Bakti Siregar

摘要

This article focuses on clustering 514 cities/districts in Indonesia based on their poverty levels from the pre-COVID-19 period (2019) until the near-end of the COVID-19 period (2022). Using publicly available data from the Central Statistics Agency (BPS) for 2019–2022, encompassing 11 poverty-related variables, the study employs unsupervised learning techniques. The process involves testing for multicollinearity; normalizing the data and applying Principal Component Analysis to reduce the variables to 2 principal components; determining the optimal number of clusters (3) is done using the elbow method, silhouette index, and gap statistic algorithm; the K-Means algorithm is then applied to cluster the cities/districts each year. The analysis reveals three clusters: cluster 1 represents areas with low poverty levels, cluster 2 comprises regions with moderate levels, and cluster 3 includes areas with high poverty rates. The number of cities/districts in cluster 3 increased from 36 in 2019 to 40 in 2022, with 32 consistently belonging to this cluster. These findings highlight the importance of targeted government assistance to these areas.