A group of spatial features whose instances are neighbors and frequently located together is called prevalent co-location pattern (PCP). In PCP mining, the task of determining the neighborhoods is crucial. Various neighborhood construction algorithms with different approaches have been proposed, however, parameter sensitivity and outlier handling are still significant issues for many of them such as distance threshold-based, density-based, k-nearest neighbors, etc. Therefore, we propose a method in order to address this problem. Firstly, an algorithm with a combination of density and connectivity called neighborhood construction (NC) is used. The process ended with a complete set of final neighborhoods for each data point (spatial instance). Subsequently, these unique neighbor sets of instances are fed into Joinless that is a traditional PCP mining algorithm based on user-predefined distance threshold to mine PCPs and we call this approach NC-j. By avoiding the usage of distance threshold in calculating neighborhood sets of each spatial instances, our method is parameter-free. Comparisons of the results mined by the proposed method and others, e.g., user-predefined distance threshold- and density-based methods, on different real-world datasets are made to demonstrate that NC-j yields better performance regarding time execution, total memory usage and scalability while maintaining high accuracy.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

An Algorithm for Discovering Prevalent Co-location Patterns with Considering Both Density and Connectivity

  • Dinhsontung Ta,
  • Phan Ha,
  • Vanha Tran,
  • Vanhieu Bui

摘要

A group of spatial features whose instances are neighbors and frequently located together is called prevalent co-location pattern (PCP). In PCP mining, the task of determining the neighborhoods is crucial. Various neighborhood construction algorithms with different approaches have been proposed, however, parameter sensitivity and outlier handling are still significant issues for many of them such as distance threshold-based, density-based, k-nearest neighbors, etc. Therefore, we propose a method in order to address this problem. Firstly, an algorithm with a combination of density and connectivity called neighborhood construction (NC) is used. The process ended with a complete set of final neighborhoods for each data point (spatial instance). Subsequently, these unique neighbor sets of instances are fed into Joinless that is a traditional PCP mining algorithm based on user-predefined distance threshold to mine PCPs and we call this approach NC-j. By avoiding the usage of distance threshold in calculating neighborhood sets of each spatial instances, our method is parameter-free. Comparisons of the results mined by the proposed method and others, e.g., user-predefined distance threshold- and density-based methods, on different real-world datasets are made to demonstrate that NC-j yields better performance regarding time execution, total memory usage and scalability while maintaining high accuracy.