<p>Large language models (LLMs) have achieved superior performance in powering text-based AI agents, endowing them with decision-making and reasoning abilities that are analogous to those exhibited by humans. Concurrently, an emerging research trend is focused on extending these LLM-powered AI agents into the <i>multimodal</i> domain. This extension facilitates the interpretation and response of AI agents to diverse multimodal user queries, thereby handling more intricate and nuanced tasks. In this paper, we conduct a systematic review of LLM-driven multimodal agents, which we refer to as <i>large multimodal agents</i> (<Emphasis FontCategory="NonProportional">LMAs</Emphasis> for short). First, we introduce the essential components involved in developing <Emphasis FontCategory="NonProportional">LMAs</Emphasis> and categorize the current body of research into four distinct types. Subsequently, we review the collaborative frameworks that integrate multiple <Emphasis FontCategory="NonProportional">LMAs</Emphasis>, with the aim of enhancing collective efficacy. One of the critical challenges in this field is the diverse evaluation methods used across existing studies, which impedes effective comparison among different <Emphasis FontCategory="NonProportional">LMAs</Emphasis>. Therefore, we compile these evaluation methodologies and establish a comprehensive framework to bridge the gaps. This framework aims to standardize evaluations, facilitating more meaningful comparisons. Concluding our review, we highlight the extensive applications of <Emphasis FontCategory="NonProportional">LMAs</Emphasis> and propose potential future research directions. Our discussion aims to provide valuable insights and guidelines for future research in this rapidly evolving field.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Large multimodal agents: a survey

  • Junlin Xie,
  • Zhihong Chen,
  • Ruifei Zhang,
  • Guanbin Li

摘要

Large language models (LLMs) have achieved superior performance in powering text-based AI agents, endowing them with decision-making and reasoning abilities that are analogous to those exhibited by humans. Concurrently, an emerging research trend is focused on extending these LLM-powered AI agents into the multimodal domain. This extension facilitates the interpretation and response of AI agents to diverse multimodal user queries, thereby handling more intricate and nuanced tasks. In this paper, we conduct a systematic review of LLM-driven multimodal agents, which we refer to as large multimodal agents (LMAs for short). First, we introduce the essential components involved in developing LMAs and categorize the current body of research into four distinct types. Subsequently, we review the collaborative frameworks that integrate multiple LMAs, with the aim of enhancing collective efficacy. One of the critical challenges in this field is the diverse evaluation methods used across existing studies, which impedes effective comparison among different LMAs. Therefore, we compile these evaluation methodologies and establish a comprehensive framework to bridge the gaps. This framework aims to standardize evaluations, facilitating more meaningful comparisons. Concluding our review, we highlight the extensive applications of LMAs and propose potential future research directions. Our discussion aims to provide valuable insights and guidelines for future research in this rapidly evolving field.