Data-driven methodologies and advanced machine learning models for predicting, evaluating, and recommending energy load measures and analysis in district area: Application of artificial intelligence
Bibliographic record
Abstract
Several methods for district energy modeling have been proposed as potential solutions to the challenges posed by building and district energy modeling. One promising approach involves using white box models, which rely on detailed physical equations and parameters to simulate energy flows and system dynamics accurately. These models provide high accuracy and insight into the physical processes within buildings and districts. However, white box models can be computationally intensive and often require extensive data inputs. Other approaches, such as black box, are also explored to balance accuracy with computational efficiency, especially when data availability or system complexity limits the practicality of white box models. Another method involves using black box models, including machine learning algorithms and artificial neural networks. Black box models focus on pattern recognition and statistical relationships within data rather than physical principles. They are particularly useful when extensive historical data is available, enabling them to effectively model complex, nonlinear relationships without detailed system knowledge. However, these models often lack interpretability, challenging the understanding of underlying processes and causal relationships. Despite this, black box models can offer high predictive accuracy and are computationally less intensive, making them suitable for large-scale district energy modeling where quick approximations are necessary. This study uses machine learning and deep neural network applications to predict and evaluate energy load metrics and conduct detailed district-level analyses. Specifically, it aims to develop models to forecast energy demand, assess load distribution, and identify consumption patterns across various zones. The study includes the application of regression, classification, and clustering methods alongside deep neural networks to effectively capture both linear and nonlinear energy consumption trends. By employing these techniques, the study seeks to generate accurate energy load predictions and uncover insights into the factors driving energy demand in complex urban environments. The research is applied to case studies involving the University of British Columbia (UBC) campus in Vancouver and the city of Stockholm in Sweden. These areas present unique energy use characteristics due to differing climates, infrastructure, and population densities, making them ideal for comparative analysis. The study’s objectives include enhancing the efficiency of energy distribution systems, identifying peak demand periods, and supporting sustainable energy planning for large district areas. Additionally, the insights gathered can inform district wide energy management practices and facilitate the integration of renewable energy sources, contributing to the development of smarter, more resilient urban energy systems. This thesis develops and evaluates data-driven methodologies and advanced machine learning (ML) models for predicting, analyzing, and recommending energy load measures in urban district energy systems. It compares white box, black box, and hybrid modeling tools, with emphasis on their applicability in energy consulting practices. The research focuses on two major case studies: the University of British Columbia (UBC) campus and the city of Stockholm. For UBC, various ML and deep learning models, including decision trees, support vector regression, and artificial neural networks, were applied to forecast electrical, hot water, and gas consumption. These models achieved high predictive accuracy (e.g., R² up to 0.94, low MAE), and their performance was validated using a structured train-test-validate split. For Stockholm, spatial clustering methods (K-means, agglomerative clustering) were used to map heat demand and allocate residual heat sources. The levelized cost of heat (LCOH) was calculated per cluster to support infrastructure planning. The results demonstrate the effectiveness of data-driven modeling for both temporal forecasting and spatial energy optimization. The proposed methodologies are generalizable and can inform scalable, lowcarbon district energy strategies.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.005 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.003 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.002 | 0.002 |
| Open science | 0.002 | 0.001 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".