<p>In real-world graph fraud detection, fraud finding often faces three limits at the same time: very few positives, shifts across domains, and tight information boundaries; If an offline test allows post-hoc (a posteriori) label rebalancing or cross-boundary information use, the offline result can deviate from what we can deploy. Many existing methods rely on larger graph models or post-processing after training, but they often depend on sampling to rebalance classes, sweeping thresholds after testing, and using loose information boundaries, so they can report numbers that are too high and are hard to reproduce. In this work, we target stable and transferable detection while keeping the real limits unchanged. We introduce <b>G-TOD</b>, a two-stage framework. In Stage&#xa0;1, we use adaptive Mixup in feature space to align source and target. In Stage&#xa0;2, we start from the Stage&#xa0;1 features and apply distribution calibration with generative oversampling for the rare class, which reduces class imbalance and helps generalization. To reduce evaluation bias, we use one strict offline setup that forbids post-hoc thresholding, keeps training and testing inside closed-domain subgraphs, and controls how many positive samples are exposed. This setup does not change the method, but it supports reproducibility and transfer under extreme imbalance and domain isolation. On several cross-domain graph datasets, <b>G-TOD</b> shows steady gains and lower variance under the same rules, which points to robustness and practical value under realistic limits.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Cross-domain graph fraud detection under extreme imbalance: a two-stage transfer framework

  • Junquan Gu,
  • Hang Yu,
  • Zhengyang Liu,
  • Dian Huang,
  • Xiangfeng Luo

摘要

In real-world graph fraud detection, fraud finding often faces three limits at the same time: very few positives, shifts across domains, and tight information boundaries; If an offline test allows post-hoc (a posteriori) label rebalancing or cross-boundary information use, the offline result can deviate from what we can deploy. Many existing methods rely on larger graph models or post-processing after training, but they often depend on sampling to rebalance classes, sweeping thresholds after testing, and using loose information boundaries, so they can report numbers that are too high and are hard to reproduce. In this work, we target stable and transferable detection while keeping the real limits unchanged. We introduce G-TOD, a two-stage framework. In Stage 1, we use adaptive Mixup in feature space to align source and target. In Stage 2, we start from the Stage 1 features and apply distribution calibration with generative oversampling for the rare class, which reduces class imbalance and helps generalization. To reduce evaluation bias, we use one strict offline setup that forbids post-hoc thresholding, keeps training and testing inside closed-domain subgraphs, and controls how many positive samples are exposed. This setup does not change the method, but it supports reproducibility and transfer under extreme imbalance and domain isolation. On several cross-domain graph datasets, G-TOD shows steady gains and lower variance under the same rules, which points to robustness and practical value under realistic limits.