In previous studies, datasets like HumanEval and MbPP were used to provide a comparative value for testing the ability of code generation with LLMs [4, 50]. These datasets often consist of function signatures, functional comments, and code snippets of the implementation. When using LLMs in the industry, these values help to find the best LLM, but are not able to provide real-world usage potentials and limitations [11].

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Related Work

  • Mahja Sarschar

摘要

In previous studies, datasets like HumanEval and MbPP were used to provide a comparative value for testing the ability of code generation with LLMs [4, 50]. These datasets often consist of function signatures, functional comments, and code snippets of the implementation. When using LLMs in the industry, these values help to find the best LLM, but are not able to provide real-world usage potentials and limitations [11].