Node classification tasks have seen considerable progress with the use of Graph Neural Networks (GNNs). However, they cannot work well on imbalanced node classification and tend to prioritize the majority classes with more labeled instances while overlooking the minority classes with fewer labeled instances. Existing solutions focus on generating new nodes to augment the training set, which may disrupt the original topological structure of the graph, so GNNs may not achieve optimal classification results. To address this issue, we introduce GraphMMC, a reliable and flexible strategy to generate pseudo-labels that can be easily integrated with various GNNs, which will augment the imbalance training set to a class-balanced set without generating new nodes. We use the similarity between the unlabeled nodes and the minority classes to correction the low-confidence pseudo-labels generated by GNNs to obtain reliable pseudo-labels. Our experiments demonstrate that the proposed method outperforms state-of-the-art baselines on several class-imbalanced datasets.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

GraphMMC: Class-Balanced Pseudo-Labels Generation for Graph Node Classification

  • Jiyou Ma,
  • Fei Chen,
  • Fan Jiang,
  • Hang Cheng,
  • Meiqing Wang

摘要

Node classification tasks have seen considerable progress with the use of Graph Neural Networks (GNNs). However, they cannot work well on imbalanced node classification and tend to prioritize the majority classes with more labeled instances while overlooking the minority classes with fewer labeled instances. Existing solutions focus on generating new nodes to augment the training set, which may disrupt the original topological structure of the graph, so GNNs may not achieve optimal classification results. To address this issue, we introduce GraphMMC, a reliable and flexible strategy to generate pseudo-labels that can be easily integrated with various GNNs, which will augment the imbalance training set to a class-balanced set without generating new nodes. We use the similarity between the unlabeled nodes and the minority classes to correction the low-confidence pseudo-labels generated by GNNs to obtain reliable pseudo-labels. Our experiments demonstrate that the proposed method outperforms state-of-the-art baselines on several class-imbalanced datasets.