Furniture assembly is a creative task, yet varying styles in assembly instructions could make them difficult to follow. In this study, we explore the potential of increasing comprehension by automatically generating text descriptions for illustration-only furniture assembly manuals using a multi-modal large-language model such as a generative pre-trained transformer (GPT). Specifically, we inputted assembly illustrations into GPT-4 and generated captions for each assembly step. These captions were added to the instructions to accommodate users who prefer textual information. We conducted two user studies to evaluate the preferences for styles of manuals and the effectiveness of the additional captions generated by GPT-4. The first study revealed that the preferable length of added captions were those less than 150 words. We also learned that captions with some errors were negligible for users. Consequently, we conducted a second study in form of a comprehension test to determine whether captions with errors are still comprehensible. The results suggest that there is no significant difference in comprehension between correct and error-containing captions generated by the model.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Additional Captions Generated by GPT-4 for Furniture Assembly Manuals

  • Ryuki Maeda,
  • Maria Larsson,
  • Hironori Yoshida

摘要

Furniture assembly is a creative task, yet varying styles in assembly instructions could make them difficult to follow. In this study, we explore the potential of increasing comprehension by automatically generating text descriptions for illustration-only furniture assembly manuals using a multi-modal large-language model such as a generative pre-trained transformer (GPT). Specifically, we inputted assembly illustrations into GPT-4 and generated captions for each assembly step. These captions were added to the instructions to accommodate users who prefer textual information. We conducted two user studies to evaluate the preferences for styles of manuals and the effectiveness of the additional captions generated by GPT-4. The first study revealed that the preferable length of added captions were those less than 150 words. We also learned that captions with some errors were negligible for users. Consequently, we conducted a second study in form of a comprehension test to determine whether captions with errors are still comprehensible. The results suggest that there is no significant difference in comprehension between correct and error-containing captions generated by the model.