This paper investigates the vision capabilities of multimodal Generative Pre-trained Transformers (GPTs) to auto-generate structured process models from diagram- and text-based documents. We introduce a dataset of 123 process models and corresponding documentation, emphasizing real-world element distributions. Using evaluation metrics for process model similarity, this enables ground truth-based assessment of process model generation. We evaluate commercial GPT capabilities with zero-, one-, and few-shot prompting strategies. Our results indicate that generative vision models can be useful tools for semi-automated process modeling based on multimodal documents. More importantly, the dataset and evaluation metrics as well as the open-source evaluation code provide a structured framework for continued systematic evaluations moving forward.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Leveraging Generative Vision Models for Extracting Process Models from Documents

  • Marvin Voelter,
  • Raheleh Hadian,
  • Timotheus Kampik,
  • Marius Breitmayer,
  • Manfred Reichert

摘要

This paper investigates the vision capabilities of multimodal Generative Pre-trained Transformers (GPTs) to auto-generate structured process models from diagram- and text-based documents. We introduce a dataset of 123 process models and corresponding documentation, emphasizing real-world element distributions. Using evaluation metrics for process model similarity, this enables ground truth-based assessment of process model generation. We evaluate commercial GPT capabilities with zero-, one-, and few-shot prompting strategies. Our results indicate that generative vision models can be useful tools for semi-automated process modeling based on multimodal documents. More importantly, the dataset and evaluation metrics as well as the open-source evaluation code provide a structured framework for continued systematic evaluations moving forward.