A high utility co-location pattern (HUCP) refers to a set of spatial features whose instances satisfy a neighbor relationship in space and have high utility values. Currently, the utility value of a co-location pattern is measured by a pattern utility ratio (PUR) metric, which is the ratio of the sum of the utility values of the features participating in the pattern to the utility of the entire data set. However, this measurement method also takes into account these features that do not participate in the pattern. In addition, the utility values of features in data sets vary greatly, when they are directly used, the utility of patterns will result in numerical deviations, which will eventually lead to missing or mining unreasonable patterns. This work first normalizes the utility values of features and then proposes a new definition to measure the utility of a pattern, i.e., pattern participation utility index (PUI), to avoid the above problems. Like the existing PUR, the newly proposed PUI also does not satisfy the downward closure property, which means that many candidates are generated during the mining process because they are not effectively pruned. This work also designs a mining algorithm based on a query style to avoid this issue. The neighbor relationships between instances are first enumerated only one times by maximal cliques and then compressed into a hash table structure. The hash table keys correspond to the initial co-location pattern candidate set. For each candidate, its co-location instances are collected by executing a query on the hash table, thus the utility of the candidate is computed quickly. The proposed method is experimentally verified on synthetic and real-world data sets. The experimental results show that the proposed method is effective and efficient compared with existing methods.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A New Method of Mining High Utility Co-location Patterns from Spatial Data

  • Vanha Tran,
  • Thiloan Bui,
  • Thaigiang Do,
  • Hoangan Le

摘要

A high utility co-location pattern (HUCP) refers to a set of spatial features whose instances satisfy a neighbor relationship in space and have high utility values. Currently, the utility value of a co-location pattern is measured by a pattern utility ratio (PUR) metric, which is the ratio of the sum of the utility values of the features participating in the pattern to the utility of the entire data set. However, this measurement method also takes into account these features that do not participate in the pattern. In addition, the utility values of features in data sets vary greatly, when they are directly used, the utility of patterns will result in numerical deviations, which will eventually lead to missing or mining unreasonable patterns. This work first normalizes the utility values of features and then proposes a new definition to measure the utility of a pattern, i.e., pattern participation utility index (PUI), to avoid the above problems. Like the existing PUR, the newly proposed PUI also does not satisfy the downward closure property, which means that many candidates are generated during the mining process because they are not effectively pruned. This work also designs a mining algorithm based on a query style to avoid this issue. The neighbor relationships between instances are first enumerated only one times by maximal cliques and then compressed into a hash table structure. The hash table keys correspond to the initial co-location pattern candidate set. For each candidate, its co-location instances are collected by executing a query on the hash table, thus the utility of the candidate is computed quickly. The proposed method is experimentally verified on synthetic and real-world data sets. The experimental results show that the proposed method is effective and efficient compared with existing methods.