Outdoor inspection and maintenance of electrical facilities are critical issues in the operation of power transmission and distribution systems. Introducing a visual dialogue interactive auxiliary system can significantly improve the work efficiency of power inspectors with less experience and reduce the risk of work errors. However, current research in the field of power inspection mainly focuses on detection models of the single modality of vision. Although it has good performance in identifying faults within the training data set, it is difficult to generate valuable responses to open-set electrical fault vocabulary. At the same time, since these models do not have interactive capabilities, besides identifying whether a fault has occurred, they can't provide any other effective auxiliary suggestions to inspectors. This paper constructs a 10K scale visual instruction-following dataset in the field of electric inspection in a cost-effective way based on the open-source image dataset InsPLAD. By using the data set constructed in this paper, the large-scale multimodal model LLaVA-v1.5 is fine-tuned to create an electrical inspection large language and vision assistant (EI-LLaVA). EI-LLaVA achieves high accuracy in traditional visual recognition tasks in the field of electrical inspection while being capable of finely describing fault types and providing maintenance suggestions.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

EI-LLaVA: Harnessing the Power of a Visual Language Dataset for Innovations in Intelligent Electrical Inspection Applications

  • Haoran Xu,
  • Muran Liu,
  • Haibin Yao,
  • Jun Meng

摘要

Outdoor inspection and maintenance of electrical facilities are critical issues in the operation of power transmission and distribution systems. Introducing a visual dialogue interactive auxiliary system can significantly improve the work efficiency of power inspectors with less experience and reduce the risk of work errors. However, current research in the field of power inspection mainly focuses on detection models of the single modality of vision. Although it has good performance in identifying faults within the training data set, it is difficult to generate valuable responses to open-set electrical fault vocabulary. At the same time, since these models do not have interactive capabilities, besides identifying whether a fault has occurred, they can't provide any other effective auxiliary suggestions to inspectors. This paper constructs a 10K scale visual instruction-following dataset in the field of electric inspection in a cost-effective way based on the open-source image dataset InsPLAD. By using the data set constructed in this paper, the large-scale multimodal model LLaVA-v1.5 is fine-tuned to create an electrical inspection large language and vision assistant (EI-LLaVA). EI-LLaVA achieves high accuracy in traditional visual recognition tasks in the field of electrical inspection while being capable of finely describing fault types and providing maintenance suggestions.