To increase the efficiency and quality of design and construction tasks, the use of Artificial Intelligence (AI) and Machine Learning (ML) offers a way to automate both repetitive and complex tasks. Many of these ML models rely heavily on large amounts of suitable, machine-readable, and labeled training data. Therefore, a variety of conceivable use cases for ML in the Architecture, Engineering and Construction (AEC) industry are difficult to implement due to a lack of freely and directly usable training data. The process of manually structuring and labeling existing data is time-consuming and needs in some cases skilled personnel to ensure the quality of the labeled data. Due to these factors, approaches for utilizing artificially generated data, referred to as synthetic data, are becoming more prevalent. Since Building Information Models contain a large amount of information, deriving training data from these models presents an obvious route for generation of this data. There are many ML applications whose implementation is inhibited due to a lack of training data, for which model-based synthetic data offer a possible solution approach. The Industry Foundation Classes (IFC) standard provides a powerful exchange format for models independently of their authoring software. Parametric and generative approaches to model creation enable the generation of numerous different building models within a short period of time and with low effort. This paper presents a workflow for automated derivation of synthetic training data from rule-based or parametrically generated models combined with existing IFC datasets as a multimodal data repository. The method is validated by testing automated synthetically labeled image data for a plan detection task, which is carried out with the Object Detection Framework YOLOv8. The suggested workflow has the potential to enhance data accessibility, thereby contributing to the implementation of ML applications in the AEC industry.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Exploring the Potential of BIM Models for Deriving Synthetic Training Data for Machine Learning Applications

  • Simon K. Hoeng,
  • Friedrich Eder,
  • Marc Schmailzl,
  • Mathias Obergrießer

摘要

To increase the efficiency and quality of design and construction tasks, the use of Artificial Intelligence (AI) and Machine Learning (ML) offers a way to automate both repetitive and complex tasks. Many of these ML models rely heavily on large amounts of suitable, machine-readable, and labeled training data. Therefore, a variety of conceivable use cases for ML in the Architecture, Engineering and Construction (AEC) industry are difficult to implement due to a lack of freely and directly usable training data. The process of manually structuring and labeling existing data is time-consuming and needs in some cases skilled personnel to ensure the quality of the labeled data. Due to these factors, approaches for utilizing artificially generated data, referred to as synthetic data, are becoming more prevalent. Since Building Information Models contain a large amount of information, deriving training data from these models presents an obvious route for generation of this data. There are many ML applications whose implementation is inhibited due to a lack of training data, for which model-based synthetic data offer a possible solution approach. The Industry Foundation Classes (IFC) standard provides a powerful exchange format for models independently of their authoring software. Parametric and generative approaches to model creation enable the generation of numerous different building models within a short period of time and with low effort. This paper presents a workflow for automated derivation of synthetic training data from rule-based or parametrically generated models combined with existing IFC datasets as a multimodal data repository. The method is validated by testing automated synthetically labeled image data for a plan detection task, which is carried out with the Object Detection Framework YOLOv8. The suggested workflow has the potential to enhance data accessibility, thereby contributing to the implementation of ML applications in the AEC industry.