The ability to automatically identify similar code fragments within huge code repositories is crucial for software development and maintenance tasks such as code reuse and debugging. Although several solutions already exist to face this challenge, not many comparisons have yet been established. For this reason, this study presents a comparative analysis of existing and emerging techniques for code similarity search. We benchmark these methods across diverse codebases, examining metrics such as indexing time, search speed, and the semantic relevance of retrieved code fragments. Our research aims to provide software developers with practical information for performing efficient code similarity searches, addressing the challenges associated with the increasing size of codebases.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Evaluation of Code Similarity Search Strategies in Large-Scale Codebases

  • Jorge Martinez-Gil,
  • Shaoyi Yin

摘要

The ability to automatically identify similar code fragments within huge code repositories is crucial for software development and maintenance tasks such as code reuse and debugging. Although several solutions already exist to face this challenge, not many comparisons have yet been established. For this reason, this study presents a comparative analysis of existing and emerging techniques for code similarity search. We benchmark these methods across diverse codebases, examining metrics such as indexing time, search speed, and the semantic relevance of retrieved code fragments. Our research aims to provide software developers with practical information for performing efficient code similarity searches, addressing the challenges associated with the increasing size of codebases.