Large language models (LLMs) constitute a breakthrough in state of-the-art Artificial Intelligence technology, which is rapidly evolving and being utilized in various domains. Applications augmented by LLMs can generate and edit images, create text based on specific prompts or assigned tasks, and extract and discuss features of images. However, given the vast amount of training data used to engineer and fine-tune these models, the output can often be misleading and baseless, a process known as hallucination. In this tutorial, the architecture of these models will be analyzed comprehensively. Evaluation paradigms and analytical tools for measuring the performance and domain-specific capacity of LLMs will be presented via a series of use cases, utilizing and showcasing tools for Image-Metadata-Analysis, Named Entity-Recognition, Knowledge-Graphs, and Multi-Genre Natural Language Inference [1–30].

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Evaluating Multimodal Large Language Models: Benchmarks, Methods, and Analytical Approaches

  • Dimitrios P. Panagoulias

摘要

Large language models (LLMs) constitute a breakthrough in state of-the-art Artificial Intelligence technology, which is rapidly evolving and being utilized in various domains. Applications augmented by LLMs can generate and edit images, create text based on specific prompts or assigned tasks, and extract and discuss features of images. However, given the vast amount of training data used to engineer and fine-tune these models, the output can often be misleading and baseless, a process known as hallucination. In this tutorial, the architecture of these models will be analyzed comprehensively. Evaluation paradigms and analytical tools for measuring the performance and domain-specific capacity of LLMs will be presented via a series of use cases, utilizing and showcasing tools for Image-Metadata-Analysis, Named Entity-Recognition, Knowledge-Graphs, and Multi-Genre Natural Language Inference [1–30].