<p>The rising energy demand of the building sector, coupled with the need for high indoor environmental quality (IEQ), necessitates intelligent control strategies. However, traditional methods like rule-based control (RBC) lack adaptability, while model predictive control (MPC) suffers from model dependency. Hence, this study designed and implemented an intelligent indoor environmental control system based on multi-sensor fusion and Building Information Modeling (BIM). At the algorithmic level, this study employed a deep fusion actor-critic reinforcement learning (DACRL) algorithm to achieve energy-efficient dynamic environmental control. The system employs a four-layer architecture with cloud-edge-end collaboration. Layer 1: The perception layer utilizes a multi-modal sensor network built on an STM32L0 low-power micro-controller and an SX1276 LoRa communication module. This architecture enables distributed collection and pre-processing of parameters such as temperature, humidity, CO<sub>2</sub>, light, and occupancy. Layer 2: The fusion layer proposes a spatio-temporal synchronization mechanism for multi-source asynchronous data based on sliding window dynamic weighting and an extended Kalman filter (EKF). As an optimization, an isolation forest and one-class support vector machine (SVM) are combined to construct a multi-layer anomaly detection process. Layer 3: The decision layer establishes a deep reinforcement learning control model based on the topology of BIM semantic mapping. The key novelty of this model lies in its innovative embedding of the Lieb–Thirring (L–T) inequality, which describes the energy concentration characteristics of spatial fields, as a spectral constraint into both the reward function and the online action projection process of the Actor-Critic algorithm. This model innovatively embeds the Lieb–Thirring (L–T) inequality, which describes the energy concentration characteristics of spatial fields, as a spectral constraint into the reward function and online action projection process of the Actor-Critic algorithm. Layer 4: The execution layer implements real-time closed-loop control of terminal equipment such as HVAC, lighting, and shading through the standard BMS protocol. This study designed large-scale simulation experiments covering six core scenarios: offices, classrooms, residences, hospitals, shopping malls, and data centers. The experimental results demonstrate that the system’s annual energy consumption per unit area decreases by up to 63.2% compared to traditional rule-based control (RBC) strategies and by 23.5% compared to model predictive control (MPC) strategies. At the same time, several indoor comfort metrics improved significantly, with CO<sub>2</sub> concentrations falling below 800 ppm 91.5% of the time and the thermal comfort (PMV) index remaining within ± 0.3 94.2% of the time. The proposed DACRL algorithm demonstrates excellent convergence and generalization capabilities. It achieves stable convergence numerically with an average of only 15,000 training steps. Performance loss in cross-building type transfer testing is only 4.9%.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A sensor-fused BIM-based ıntelligent control system for energy-efficient ındoor environmental regulation using deep actor-critic reinforcement learning (DACRL)

  • Libin Tong

摘要

The rising energy demand of the building sector, coupled with the need for high indoor environmental quality (IEQ), necessitates intelligent control strategies. However, traditional methods like rule-based control (RBC) lack adaptability, while model predictive control (MPC) suffers from model dependency. Hence, this study designed and implemented an intelligent indoor environmental control system based on multi-sensor fusion and Building Information Modeling (BIM). At the algorithmic level, this study employed a deep fusion actor-critic reinforcement learning (DACRL) algorithm to achieve energy-efficient dynamic environmental control. The system employs a four-layer architecture with cloud-edge-end collaboration. Layer 1: The perception layer utilizes a multi-modal sensor network built on an STM32L0 low-power micro-controller and an SX1276 LoRa communication module. This architecture enables distributed collection and pre-processing of parameters such as temperature, humidity, CO2, light, and occupancy. Layer 2: The fusion layer proposes a spatio-temporal synchronization mechanism for multi-source asynchronous data based on sliding window dynamic weighting and an extended Kalman filter (EKF). As an optimization, an isolation forest and one-class support vector machine (SVM) are combined to construct a multi-layer anomaly detection process. Layer 3: The decision layer establishes a deep reinforcement learning control model based on the topology of BIM semantic mapping. The key novelty of this model lies in its innovative embedding of the Lieb–Thirring (L–T) inequality, which describes the energy concentration characteristics of spatial fields, as a spectral constraint into both the reward function and the online action projection process of the Actor-Critic algorithm. This model innovatively embeds the Lieb–Thirring (L–T) inequality, which describes the energy concentration characteristics of spatial fields, as a spectral constraint into the reward function and online action projection process of the Actor-Critic algorithm. Layer 4: The execution layer implements real-time closed-loop control of terminal equipment such as HVAC, lighting, and shading through the standard BMS protocol. This study designed large-scale simulation experiments covering six core scenarios: offices, classrooms, residences, hospitals, shopping malls, and data centers. The experimental results demonstrate that the system’s annual energy consumption per unit area decreases by up to 63.2% compared to traditional rule-based control (RBC) strategies and by 23.5% compared to model predictive control (MPC) strategies. At the same time, several indoor comfort metrics improved significantly, with CO2 concentrations falling below 800 ppm 91.5% of the time and the thermal comfort (PMV) index remaining within ± 0.3 94.2% of the time. The proposed DACRL algorithm demonstrates excellent convergence and generalization capabilities. It achieves stable convergence numerically with an average of only 15,000 training steps. Performance loss in cross-building type transfer testing is only 4.9%.