This paper addresses the underexplored evaluation of Retrieval-Augmented Generation (RAG) systems in the financial sector, specifically focusing on mergers and acquisitions (M&A). We developed a comprehensive assessment framework to evaluate existing RAG solutions for M&A companies. Engaging financial experts, we compiled a dataset of frequently asked questions and benchmark answers. Multiple RAG systems—including Azure GPT-4, AWS Q, and others—were configured to generate answers to these questions. An evaluation application was implemented to automatically assess the quality of the generated answers using thirteen key performance indicators, combining large language models and human expert evaluations. Consolidated into a standardized leaderboard, the results revealed that Azure GPT-4 with Azure AI Search outperformed other systems by a 15% margin, demonstrating high accuracy and relevance. Our study provides organizations with actionable insights for selecting appropriate RAG solutions, highlights areas for improvement in existing systems, and contributes to advancing RAG technologies in the financial industry.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Analysis of the Efficiency and Comparison of Retrieval-Augmented Generation Systems in Mergers and Acquisitions: Recommendations for Choosing the Best Solutions

  • Oleg Barabash,
  • Danylo Vorvul,
  • Andrii Musienko,
  • Vladyslav Bezsmertniy

摘要

This paper addresses the underexplored evaluation of Retrieval-Augmented Generation (RAG) systems in the financial sector, specifically focusing on mergers and acquisitions (M&A). We developed a comprehensive assessment framework to evaluate existing RAG solutions for M&A companies. Engaging financial experts, we compiled a dataset of frequently asked questions and benchmark answers. Multiple RAG systems—including Azure GPT-4, AWS Q, and others—were configured to generate answers to these questions. An evaluation application was implemented to automatically assess the quality of the generated answers using thirteen key performance indicators, combining large language models and human expert evaluations. Consolidated into a standardized leaderboard, the results revealed that Azure GPT-4 with Azure AI Search outperformed other systems by a 15% margin, demonstrating high accuracy and relevance. Our study provides organizations with actionable insights for selecting appropriate RAG solutions, highlights areas for improvement in existing systems, and contributes to advancing RAG technologies in the financial industry.