Graph mixup data augmentation has been introduced to enhance the generalization of Graph Neural Networks across various graph classification tasks. Existing methods interpolate graphs and their labels by fixed ratios sampled from a Beta distribution, creating synthetic graphs for training. However, these methods fail to notice the lack of node correspondence between the graphs to be mixed, resulting in structural distortion and label mismatch during the mixing process, which negatively affects the model performance. In this work, we propose SimMix, a simple yet effective framework that adaptively adjusts the label mixing ratio based on graph similarity. We incorporate a learnable similarity function to compute more accurate mixing ratios, and utilize supervised contrastive learning to further refine this process by enforcing separation in the embedding space, thereby improving label consistency. Extensive experiments conducted across seven public datasets demonstrate that SimMix achieves state-of-the-art performance in the majority of cases, particularly excelling in scenarios against noisy labels.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

SimMix: Enhancing Label Consistency in Graph Mixup for Improved Graph Classification

  • Xingkai Yao,
  • Zhongqiu Chen,
  • Zongxing Zhao,
  • Xiaofang Zhang

摘要

Graph mixup data augmentation has been introduced to enhance the generalization of Graph Neural Networks across various graph classification tasks. Existing methods interpolate graphs and their labels by fixed ratios sampled from a Beta distribution, creating synthetic graphs for training. However, these methods fail to notice the lack of node correspondence between the graphs to be mixed, resulting in structural distortion and label mismatch during the mixing process, which negatively affects the model performance. In this work, we propose SimMix, a simple yet effective framework that adaptively adjusts the label mixing ratio based on graph similarity. We incorporate a learnable similarity function to compute more accurate mixing ratios, and utilize supervised contrastive learning to further refine this process by enforcing separation in the embedding space, thereby improving label consistency. Extensive experiments conducted across seven public datasets demonstrate that SimMix achieves state-of-the-art performance in the majority of cases, particularly excelling in scenarios against noisy labels.