Outlier detection is crucial for data cleaning, influencing analysis and decision-making. While numerical outlier detection is well-studied, identifying outliers in relational data with categorical attributes poses greater challenges due to difficulties in defining a suitable similarity measure. Current approaches for detecting categorical outliers are based on coding the categorical values as numerical values, using the frequency as an indicator of the outlierness score and extracting predefined syntactic structures of the values. In this paper, we propose FIONA (FInding Outliers iN Attributes) to detect outliers in attributes with categorical values. Since categorical values in the relational model usually follow specific syntactic structures, FIONA defines a similarity measure that can reveal the hidden patterns and identify a set of dominant patterns in the data. Values that do not conform to the dominating patterns are declared as outliers. In comparison to alternative tools, FIONA accurately identifies outliers and dominant patterns within datasets and provides a clear explanation for declaring a given value as an outlier.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

FIONA: Detecting Syntactical Outliers in Attributes with Categorical Values

  • Thanos Tsiamis,
  • Abdulhakim A. Qahtan

摘要

Outlier detection is crucial for data cleaning, influencing analysis and decision-making. While numerical outlier detection is well-studied, identifying outliers in relational data with categorical attributes poses greater challenges due to difficulties in defining a suitable similarity measure. Current approaches for detecting categorical outliers are based on coding the categorical values as numerical values, using the frequency as an indicator of the outlierness score and extracting predefined syntactic structures of the values. In this paper, we propose FIONA (FInding Outliers iN Attributes) to detect outliers in attributes with categorical values. Since categorical values in the relational model usually follow specific syntactic structures, FIONA defines a similarity measure that can reveal the hidden patterns and identify a set of dominant patterns in the data. Values that do not conform to the dominating patterns are declared as outliers. In comparison to alternative tools, FIONA accurately identifies outliers and dominant patterns within datasets and provides a clear explanation for declaring a given value as an outlier.