With the continuous development of large language models (LLMs), researchers have created and introduced a significant number of various code generation models. Existing code generation evaluation benchmarks, such as HumanEval, HumanEVAL-X, MBPP, MBXP and APPS, can measure functional correctness for synthesizing programs from docstrings with general requirements. The problems in these datasets are all in English, mainly testing the code generation ability of large models when facing English questions. With the continuous development of Chinese large models, the code generation ability of Chinese large models is also worth exploring. We need a code generation ability that can test the code generation ability of large models when facing English and Chinese questions. In this paper, we evaluate 16 state-of-the-art large language models for program synthesis in terms of doctrings describing long chains of operations and complex requirements, on a new benchmark, Hard400. It contains 422 Leet- code high-difficulty level coding questions in both English and Chinese formats. Our dataset mainly consists of leetcode high-difficulty level questions.This indicates that Hard400 is more difficult and challenging.Among them, GPT shows the best code generation capabilities.However, large language models still need to be improved in terms of code generation. We hope that Hrad400 can be used as a benchmark to evaluate the code generation capabilities of large models and further promote the development and growth of large models.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Hard400: A Bilingual Code Generation Evaluation Benchmark for Large Language Models

  • Qin Zhang,
  • Song Zhang,
  • Junjie Li,
  • Hong Zhou,
  • Xiaojun Chen,
  • Han Liu

摘要

With the continuous development of large language models (LLMs), researchers have created and introduced a significant number of various code generation models. Existing code generation evaluation benchmarks, such as HumanEval, HumanEVAL-X, MBPP, MBXP and APPS, can measure functional correctness for synthesizing programs from docstrings with general requirements. The problems in these datasets are all in English, mainly testing the code generation ability of large models when facing English questions. With the continuous development of Chinese large models, the code generation ability of Chinese large models is also worth exploring. We need a code generation ability that can test the code generation ability of large models when facing English and Chinese questions. In this paper, we evaluate 16 state-of-the-art large language models for program synthesis in terms of doctrings describing long chains of operations and complex requirements, on a new benchmark, Hard400. It contains 422 Leet- code high-difficulty level coding questions in both English and Chinese formats. Our dataset mainly consists of leetcode high-difficulty level questions.This indicates that Hard400 is more difficult and challenging.Among them, GPT shows the best code generation capabilities.However, large language models still need to be improved in terms of code generation. We hope that Hrad400 can be used as a benchmark to evaluate the code generation capabilities of large models and further promote the development and growth of large models.