Large Language Models (LLMs) are a type of artificial intelligence capable of processing and generating natural language. These models are trained on vast amounts of data, such as text and code, which enables them to perform various tasks like text generation, language translation, question answering, text summarization, etc.. The purpose of this research was to find an LLM that meets the following requirements: 1)easy to implement with an understanding of the Spanish and English languages, and 2)accurately answers diagnostic questions and action plans related to Industry 4.0 (I4.0). Some open-source LLMs were selected, each with 7B parameters and quantizations of Q2 and Q4 for the application of a first set of general tests. The results showed that the model with the best performance in both languages English and Spanish was Mistral. A second set of Spanish tests was then conducted with and without Retrieval-Augmented Generation (RAG), using three documents related to the topic of I4.0. The results demonstrated Mistral’s capability with RAG to answer questions on the studied context with 95 % accuracy.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Performance Testing of Large Language Models in the Context of Industry 4.0

  • Pedro José González Díaz,
  • Ailín Orjuela Duarte,
  • William Mauricio Rojas Contreras,
  • Luz Marina Santos Jaimes

摘要

Large Language Models (LLMs) are a type of artificial intelligence capable of processing and generating natural language. These models are trained on vast amounts of data, such as text and code, which enables them to perform various tasks like text generation, language translation, question answering, text summarization, etc.. The purpose of this research was to find an LLM that meets the following requirements: 1)easy to implement with an understanding of the Spanish and English languages, and 2)accurately answers diagnostic questions and action plans related to Industry 4.0 (I4.0). Some open-source LLMs were selected, each with 7B parameters and quantizations of Q2 and Q4 for the application of a first set of general tests. The results showed that the model with the best performance in both languages English and Spanish was Mistral. A second set of Spanish tests was then conducted with and without Retrieval-Augmented Generation (RAG), using three documents related to the topic of I4.0. The results demonstrated Mistral’s capability with RAG to answer questions on the studied context with 95 % accuracy.