The Mallows model is a sampling technique frequently used to create synthetic election datasets and evaluate ranking aggregation algorithms. It provides flexibility, as authors can recreate different election scenarios by modifying the model’s parameters. When applying this model, results from data with different numbers of voters are often compared without adequate consideration, mistakenly assuming a consistent underlying structure in the preference distribution. However, experimental results have shown that this practice may be problematic because some intrinsic properties of the data fluctuate when modifying this parameter. In this paper, we address this issue by studying the behaviour of Mallows model for producing elections when increasing the number of voters. We generate a synthetic dataset that contains simulations of elections for different numbers of voters and alternatives and measure some structural characteristics of the generated data. Afterwards, we analyse the behaviour of these properties as the number of voters increases. Results show that some aspects, such as the modal ranking, are affected by this variation. This leads to the conclusion that the performance of ranking aggregation algorithms, particularly those based on the Condorcet method, cannot be directly compared to data with different numbers of voters due to varying aggregation difficulty.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Understanding Data Properties in the Mallows Model: Impact of Voter Count Variability

  • Mario Villar,
  • Noelia Rico,
  • Irene Díaz

摘要

The Mallows model is a sampling technique frequently used to create synthetic election datasets and evaluate ranking aggregation algorithms. It provides flexibility, as authors can recreate different election scenarios by modifying the model’s parameters. When applying this model, results from data with different numbers of voters are often compared without adequate consideration, mistakenly assuming a consistent underlying structure in the preference distribution. However, experimental results have shown that this practice may be problematic because some intrinsic properties of the data fluctuate when modifying this parameter. In this paper, we address this issue by studying the behaviour of Mallows model for producing elections when increasing the number of voters. We generate a synthetic dataset that contains simulations of elections for different numbers of voters and alternatives and measure some structural characteristics of the generated data. Afterwards, we analyse the behaviour of these properties as the number of voters increases. Results show that some aspects, such as the modal ranking, are affected by this variation. This leads to the conclusion that the performance of ranking aggregation algorithms, particularly those based on the Condorcet method, cannot be directly compared to data with different numbers of voters due to varying aggregation difficulty.