<p>Intelligent and autonomous networks require precise and fast mechanisms to minimize errors and ensure efficient operation. Modern methods are increasingly based on artificial intelligence, in particular on machine learning, to reliably process large amounts of data. While high-quality datasets are essential to train machine learning models, assessing the quality of datasets can be challenging and is often overlooked or underestimated. This paper proposes a novel method of permutation testing to assess one relevant dataset quality dimension in the context of binary or multiclass classification problems: the strength of the association between data and labels. The method described is called Permutation for Quality of Dataset Assessment (PerQoDA). In this paper we introduce the method, the statistics and visualisations necessary for proper interpretation of the results, and lastly, the theoretical justification of limits of performance. According to our experiments carried out on both simulated as well as real network datasets, the PerQoDA method can correctly estimate the strength of relationships in labelled datasets across a range of scenarios.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

PerQoDA: strength of association between data and labels as a measure of dataset quality in network traffic classification

  • Katarzyna Wasielewska,
  • Dominik Soukup,
  • Tomáš Čejka,
  • Joe Carthy,
  • José Camacho

摘要

Intelligent and autonomous networks require precise and fast mechanisms to minimize errors and ensure efficient operation. Modern methods are increasingly based on artificial intelligence, in particular on machine learning, to reliably process large amounts of data. While high-quality datasets are essential to train machine learning models, assessing the quality of datasets can be challenging and is often overlooked or underestimated. This paper proposes a novel method of permutation testing to assess one relevant dataset quality dimension in the context of binary or multiclass classification problems: the strength of the association between data and labels. The method described is called Permutation for Quality of Dataset Assessment (PerQoDA). In this paper we introduce the method, the statistics and visualisations necessary for proper interpretation of the results, and lastly, the theoretical justification of limits of performance. According to our experiments carried out on both simulated as well as real network datasets, the PerQoDA method can correctly estimate the strength of relationships in labelled datasets across a range of scenarios.