<p>As software evolves, groups of variables often recur, forming <i>data clumps</i>. This study extends our previous work by applying our detection tool to 16 public repositories covering 2696 tagged versions. We mined over 6&#xa0;million additional data clumps and updated the publicly available dataset. The analysis confirms that data clumps persist over time and tend to grow into larger clusters, which hinders refactoring. Variables such as <i>name</i> and <i>key</i> dominate across projects, whereas credential variables occur rarely. These results consolidate earlier observations and provide a stronger foundation for studying the connection between data clumps and faults.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Unveiling Data Clumps: A Detailed Longitudinal Analysis on Software Quality Across Public Repositories

  • Nils Baumgartner,
  • Elke Pulvermüller

摘要

As software evolves, groups of variables often recur, forming data clumps. This study extends our previous work by applying our detection tool to 16 public repositories covering 2696 tagged versions. We mined over 6 million additional data clumps and updated the publicly available dataset. The analysis confirms that data clumps persist over time and tend to grow into larger clusters, which hinders refactoring. Variables such as name and key dominate across projects, whereas credential variables occur rarely. These results consolidate earlier observations and provide a stronger foundation for studying the connection between data clumps and faults.