Fault detection and diagnosis in light commercial buildings’ HVAC systems: A comprehensive framework, application, and performance evaluation
Bibliographic record
Abstract
The data-driven approach currently dominates the field of Automatic Fault Detection and Diagnosis (AFDD) in HVAC systems. However, a significant concern lies in the prevalent use of labeled experimental and simulation data, which often does not represent real-world operational conditions. This study unveils a comprehensive framework for AFDD in light commercial buildings, effectively leveraging unlabeled raw data extracted directly from their Building BAS. Its main goal is to provide a versatile methodology tailored for real-world applicability. Buildings classified as “light commercial” typically have less than 2,500 square meters of floor area and no more than six stories, such as small offices, medical facilities, banks, small manufacturing facilities, etc. A common feature of these buildings is the fact that the HVAC systems tend to be relatively simple and have similar configurations, thus making it easy to develop scalable and reproducible fault detection methods. The study focuses on a practical case study in a light commercial building HVAC system situated in Montreal, Canada, encompassing a single Air Handling Unit (AHU) and four Variable Air Volume (VAV) reheating boxes to evaluate the framework. This comprehensive framework encompasses a sequence of sub-objectives: creating a sizable, synchronized raw dataset from diverse BAS sensor tags, comprehensive data cleansing to address inconsistencies, developing an anomaly detection method, investigating these anomalies to extract underlying rules, and finally, dataset labeling. An AFDD classification model is then applied to evaluate its ability to distinguish normal from faulty conditions across various fault types. The study highlights the potential of dimensional reduction techniques and unsupervised clustering for effective anomaly detection in light commercial buildings, as well as the power of the Decision Tree classifier for uncovering hidden patterns, especially in anomaly conditions. It also highlights the significance of addressing imbalanced datasets in AFDD and the complexities of detecting sizing-related faults. Despite these challenges, the framework exhibits robust performance in detecting and diagnosing a range of HVAC faults. It offers a systematic and adaptable approach for handling real-world operational data in light commercial building HVAC systems, extendible to other building types, bridging the gap between data-driven methods and practical applications.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".