<p>Graphs serve as fundamental data structures for representing relationships and interactions among entities. Utilizing graphs to represent data helps in uncovering complex relationships and patterns that could be missed in models concentrating on individual data points. Attributed graphs enhance this capability by associating data features with vertices or edges of the graph. However, when addressing complex real-world problems, datasets often involve numerous features. Feature selection emerges as a critical technique that aims to identify a pertinent subset of features for specific tasks such as classification, prediction, or anomaly detection. Nevertheless, the computational demands of feature selection are heightened by the size and complexity of these datasets. Furthermore, the domain of attributed graphs faces a deficiency in adequate feature selection methods, leading to suboptimal outcomes in various data analysis tasks. This study attempts to address this challenge by framing the selection of features in attributed graph data as a graph similarity problem. We evaluated the applicability of our graph-based feature selection approach through two case studies utilizing a graph constructed from Brazilian census data. The first case study focused on identifying key census features associated with <InlineEquation ID="IEq3"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="13748_2025_407_Article_IEq1.gif" Format="GIF" Height="16" Rendition="HTML" Resolution="72" Type="Linedraw" Width="32" /> </InlineMediaObject> <EquationSource Format="TEX">\(\hbox {CO}_{2}\)</EquationSource> <EquationSource Format="MATHML"><math> <msub> <mtext>CO</mtext> <mn>2</mn> </msub> </math></EquationSource> </InlineEquation> emissions in Brazil. The second case study aimed to uncover socio-economic determinants, derived from census data, of Brazilian homicide rates.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Feature selection in graph attributed data: case studies on real-world \(\hbox {CO}_{2}\) emissions and homicide rates

  • Greta Augat Abib,
  • Ronaldo Cristiano Prati

摘要

Graphs serve as fundamental data structures for representing relationships and interactions among entities. Utilizing graphs to represent data helps in uncovering complex relationships and patterns that could be missed in models concentrating on individual data points. Attributed graphs enhance this capability by associating data features with vertices or edges of the graph. However, when addressing complex real-world problems, datasets often involve numerous features. Feature selection emerges as a critical technique that aims to identify a pertinent subset of features for specific tasks such as classification, prediction, or anomaly detection. Nevertheless, the computational demands of feature selection are heightened by the size and complexity of these datasets. Furthermore, the domain of attributed graphs faces a deficiency in adequate feature selection methods, leading to suboptimal outcomes in various data analysis tasks. This study attempts to address this challenge by framing the selection of features in attributed graph data as a graph similarity problem. We evaluated the applicability of our graph-based feature selection approach through two case studies utilizing a graph constructed from Brazilian census data. The first case study focused on identifying key census features associated with \(\hbox {CO}_{2}\) CO 2 emissions in Brazil. The second case study aimed to uncover socio-economic determinants, derived from census data, of Brazilian homicide rates.