High utility co-location pattern (HUCP) mining refers to discovering a group of spatial features from spatial data which the instances of the group are neighbors in space and have high utility values. Currently, existing algorithms use a unique minimum utility threshold to filter HUCPs, but instance utility values and instance numbers of features in datasets vary greatly, thus it is not fair to compare all features with a unique threshold, which can easily filter out some interesting patterns. This paper proposes an algorithm for mining HUCPs based on multiple utility thresholds to void this issue. Users set a utility threshold for each feature according to their application field. The threshold of a pattern is adaptively determined by the features involved in it, i.e., the minimum threshold of these features that participate in the pattern. Moreover, to avoid repeatedly verifying the neighboring relationships between instances when collecting participating instances of each pattern, this paper adopts a mining method that uses maximal cliques and a hash table structure. First, neighboring instances are enumerated only once by maximal cliques and put into a close-packed hash structure. Then, the instances of features that participate in any patterns are queried from this structure without repeatedly verifying neighbor relationships. The designed algorithm is experimentally assessed on both synthetic and real-world spatial data. The results prove that the designed algorithm is more effective and efficient than existing algorithms.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A High Utility Co-location Pattern Mining Algorithm Using Multiple Utility Thresholds

  • Vanha Tran,
  • Thiloan Bui,
  • Thaigiang Do,
  • Hoangan Le

摘要

High utility co-location pattern (HUCP) mining refers to discovering a group of spatial features from spatial data which the instances of the group are neighbors in space and have high utility values. Currently, existing algorithms use a unique minimum utility threshold to filter HUCPs, but instance utility values and instance numbers of features in datasets vary greatly, thus it is not fair to compare all features with a unique threshold, which can easily filter out some interesting patterns. This paper proposes an algorithm for mining HUCPs based on multiple utility thresholds to void this issue. Users set a utility threshold for each feature according to their application field. The threshold of a pattern is adaptively determined by the features involved in it, i.e., the minimum threshold of these features that participate in the pattern. Moreover, to avoid repeatedly verifying the neighboring relationships between instances when collecting participating instances of each pattern, this paper adopts a mining method that uses maximal cliques and a hash table structure. First, neighboring instances are enumerated only once by maximal cliques and put into a close-packed hash structure. Then, the instances of features that participate in any patterns are queried from this structure without repeatedly verifying neighbor relationships. The designed algorithm is experimentally assessed on both synthetic and real-world spatial data. The results prove that the designed algorithm is more effective and efficient than existing algorithms.