Related Work
摘要
In previous studies, datasets like HumanEval and MbPP were used to provide a comparative value for testing the ability of code generation with LLMs [4, 50]. These datasets often consist of function signatures, functional comments, and code snippets of the implementation. When using LLMs in the industry, these values help to find the best LLM, but are not able to provide real-world usage potentials and limitations [11].