Using surrogate models to analyze the impact of geometry on the energy efficiency of buildings
Bibliographic record
Abstract
In recent times data-driven approaches to parametrically optimize and explore \nbuilding geometry has been proven to be a powerful tool that can replace computationally expensive and time-consuming simulations for energy prediction in the early \ndesign process. In this research, we explore the use of surrogate models, i.e. efficient \nstatistical approximations of expensive physics-based building simulation models, to \nlower the computational burden of large-scale building geometry analysis. We try \ndifferent approaches and techniques to train a machine learning model using multiple \ndatasets to analyze the impact of geometry and envelope features on the energy efficiency of buildings. These contributions are presented in the form of two conference \npapers and one journal paper (being prepared for submission) that iteratively build \nup the underlying methodology. \nThe first conference paper contains preliminary experiments using 4 manually \ngenerated building geometries for office buildings. Data were generated by simulating various building samples in EnergyPlus for different geometries. We used the \ngenerated data to train a machine learning model using support vector regression. \nWe trained two separate models for predicting heating and cooling loads. The lesson \nlearned from this first experiment was that the prediction of the models was not great \ndue to insufficient geometric features explaining the variability in geometry and the \nlack of sufficient data for varied geometries. \nThe second conference paper developed a novel dataset of 38,000 building energy \nmodels for varied geometry using 2D images of real-world residences. We developed \na workflow in the Grasshopper/Rhino environment which can convert 2D images of \na floor plan into a vector format then into a building energy model ready to be simulated in EnergyPlus. The workflow can also extract up to 20 geometric features from \nthe model, to be used as features in the machine learning process. We used these \nfeatures and the simulation results to train a neural network-based surrogate model. \nA sensitivity analysis was performed to understand the impact and importance of \neach feature to the energy use of the building. From the results of the experiment, \nwe found that off-the-shelf neural network-based surrogates provided with engineered \nfeatures can very well emulate the desired simulation outputs. We also repeated \nthe experiment for 6 different climatic zones across Canada to understand the impact of geometric features across various climates; these findings are presented in an \nappendix. \niv \nIn the journal paper, we explored two different methodologies to train surrogate \nmodels: monolithic and component-based. We explored the component-based modeling technique as it allows the model to be more versatile if we need to add more \ncomponents to it, ultimately increasing the usability of the model. We conducted \nfurther experiments by adding complexity to the geometry surrogate model. We introduced 10 envelope features as an input to the surrogate along with the 20 geometric \nfeatures. We trained 6 different surrogate models using different datasets by varying \ngeometric and envelope features. From the results of the experiment, we found that \nthe monolithic model performs the best but the component-based surrogate also falls \ninto an acceptable range of accuracy. \nFrom the overall results across the three papers, we see that simple neural network-based surrogate models perform really well to emulate simulation outcomes over a \nwide variety of geometries and envelope features
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".