This paper presents the benchmarking of three multi-agent systems powered by large language models. The paper presents a comparative analysis of AutoGen, CrewAI, and TaskWeaver. Nowadays, large language models have emerged as powerful tools able to assist users in various areas. The integration of large language models into multi-agent systems increases their potential for collaborative problem-solving. This study focuses on a case study involving a machine learning code generation task which is used to evaluate the framework’s performance. To assess the performance of the solutions, it is requested to create energy forecasting models using the same dataset as the base. After producing the code, a new dataset is used to test the model performance using the root mean square error. The three solutions were able to provide results using multiple large language models. The best result was achieved by TaskWeaver using GPT-3.5, with an error of 25.04.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Benchmarking Large Language Models for Multi-agent Systems: A Comparative Analysis of AutoGen, CrewAI, and TaskWeaver

  • Rafael Barbarroxa,
  • Luis Gomes,
  • Zita Vale

摘要

This paper presents the benchmarking of three multi-agent systems powered by large language models. The paper presents a comparative analysis of AutoGen, CrewAI, and TaskWeaver. Nowadays, large language models have emerged as powerful tools able to assist users in various areas. The integration of large language models into multi-agent systems increases their potential for collaborative problem-solving. This study focuses on a case study involving a machine learning code generation task which is used to evaluate the framework’s performance. To assess the performance of the solutions, it is requested to create energy forecasting models using the same dataset as the base. After producing the code, a new dataset is used to test the model performance using the root mean square error. The three solutions were able to provide results using multiple large language models. The best result was achieved by TaskWeaver using GPT-3.5, with an error of 25.04.