<p>Fuzzy Pattern Trees (FPTs) are symbolic tree-based structures whose internal nodes are fuzzy operators, and the leaves are fuzzy features, which enhance interpretability by representing data with meaningful fuzzy terms. However, conventional FPT approaches typically employ uniformly distributed membership functions, which often fail to accurately represent features in real-world datasets. In this work, we propose an automatic method to adapt the bounds of fuzzy features based on their data distributions, with a focus on a simple triangular membership scheme. We evaluate our approach across 11 benchmark classification problems, incorporating six parsimony pressure methods to promote more compact solutions. Our results demonstrate that the adapted fuzzification scheme, beyond improving interpretability, consistently yields models that better balance accuracy and size when compared to uniform representations, appearing on the Pareto front 20 times, while the second-best scheme appeared only 15 times.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

On fitting numerical features into probabilistic distributions to represent data for fuzzy pattern trees

  • Allan de Lima,
  • Juan F. H. Albarracín,
  • Douglas Mota Dias,
  • Jorge Amaral,
  • Conor Ryan

摘要

Fuzzy Pattern Trees (FPTs) are symbolic tree-based structures whose internal nodes are fuzzy operators, and the leaves are fuzzy features, which enhance interpretability by representing data with meaningful fuzzy terms. However, conventional FPT approaches typically employ uniformly distributed membership functions, which often fail to accurately represent features in real-world datasets. In this work, we propose an automatic method to adapt the bounds of fuzzy features based on their data distributions, with a focus on a simple triangular membership scheme. We evaluate our approach across 11 benchmark classification problems, incorporating six parsimony pressure methods to promote more compact solutions. Our results demonstrate that the adapted fuzzification scheme, beyond improving interpretability, consistently yields models that better balance accuracy and size when compared to uniform representations, appearing on the Pareto front 20 times, while the second-best scheme appeared only 15 times.