Large Language Models (LLMs) have been used in numerous studies to pass written examinations, with mixed performances. Prior studies have relied on the text capability of LLMs and close ended questions (e.g., multiple choices), often examining performances without comparison to real-world test-takers. Recent studies have emphasized the need to test multiple LLMs and situate performances with respect to students’ scores. In addition, recent LLMs can handle both text and images. In our empirical study, we leverage this technical improvement to examine whether LLMs can learn from the same material (i.e., readings and slide decks) that was given to students in a graduate class on conceptual modeling. Our experiments on GPT-4o and Claude 3.5 Sonnet demonstrate that both LLMs improve performances by learning from the material across five tests (covering different aspects of conceptual modeling) and across levels of mastery (e.g., understand, create). Both LLMs achieved passing grades. GPT-4o was often aligned with, or above, a median student. Our results pave the way for future studies in which LLMs may take a course in conceptual modeling, by decomposing video recordings of lectures into images.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Can Large Language Models Learn Conceptual Modeling by Looking at Slide Decks and Pass Graduate Examinations? An Empirical Study

  • Noé Y. Flandre,
  • Philippe J. Giabbanelli

摘要

Large Language Models (LLMs) have been used in numerous studies to pass written examinations, with mixed performances. Prior studies have relied on the text capability of LLMs and close ended questions (e.g., multiple choices), often examining performances without comparison to real-world test-takers. Recent studies have emphasized the need to test multiple LLMs and situate performances with respect to students’ scores. In addition, recent LLMs can handle both text and images. In our empirical study, we leverage this technical improvement to examine whether LLMs can learn from the same material (i.e., readings and slide decks) that was given to students in a graduate class on conceptual modeling. Our experiments on GPT-4o and Claude 3.5 Sonnet demonstrate that both LLMs improve performances by learning from the material across five tests (covering different aspects of conceptual modeling) and across levels of mastery (e.g., understand, create). Both LLMs achieved passing grades. GPT-4o was often aligned with, or above, a median student. Our results pave the way for future studies in which LLMs may take a course in conceptual modeling, by decomposing video recordings of lectures into images.