We demonstrate the use of principal component analysis (PCA) and exploratory factor analysis (EFA) to conduct spatio-temporal analysis of a crime dataset. We use a 30000-records crime dataset for the city of Los Angeles for the year 2020 as the dataset for our analysis. We extract a matrix (T-L) of the time (hourly index: 0, 1, 2,…, 23) vs. location type (a total of 28 location types) whose entries store the values for the number of crimes reported at a particular location type during the time spanning an hour of the day, across the entire dataset. We conduct PCA on the T-L matrix and its transpose to respectively rank and classify hours of the day and the location types to one of the three categories: hotspots, tepidspots and coldspots. We use the weighted average scores of the principal components and the properties of the principal components to facilitate such quantification, ranking and classification. By conducting EFA on the T-L matrix, we observe clusters of location types that could have a common latent factor/theme (like location types in the vicinity of a residence; location types involving dwelling units/inside a residence; commercial location types such as stores/businesses, etc.) so that we could categorize clusters of location types as “hotspots” or “not hotspots”.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Spatio-Temporal Analysis of a Crime Dataset Using Principal Component Analysis and Exploratory Factor Analysis

  • Natarajan Meghanathan

摘要

We demonstrate the use of principal component analysis (PCA) and exploratory factor analysis (EFA) to conduct spatio-temporal analysis of a crime dataset. We use a 30000-records crime dataset for the city of Los Angeles for the year 2020 as the dataset for our analysis. We extract a matrix (T-L) of the time (hourly index: 0, 1, 2,…, 23) vs. location type (a total of 28 location types) whose entries store the values for the number of crimes reported at a particular location type during the time spanning an hour of the day, across the entire dataset. We conduct PCA on the T-L matrix and its transpose to respectively rank and classify hours of the day and the location types to one of the three categories: hotspots, tepidspots and coldspots. We use the weighted average scores of the principal components and the properties of the principal components to facilitate such quantification, ranking and classification. By conducting EFA on the T-L matrix, we observe clusters of location types that could have a common latent factor/theme (like location types in the vicinity of a residence; location types involving dwelling units/inside a residence; commercial location types such as stores/businesses, etc.) so that we could categorize clusters of location types as “hotspots” or “not hotspots”.