Advancements in supervised machine learning for outdoor thermal comfort: A comprehensive systematic review of scales, applications, and data types
Bibliographic record
Abstract
• Systematic review of articles on the use of supervised machine learning (SML) for outdoor thermal environment (OTC) research. • The application of SML algorithms in OTC was discussed to provide reference and understanding for a wider audience. • The limitations, opportunities, and significance of the current application of SML in OTC research have been identified. Supervised Machine Learning (SML) has a proven track record in addressing the complexities of urban studies. To enhance the accuracy and efficiency of current urban outdoor thermal environment research, many scholars had employed SML methods. However, the current status and limitations of SML applications in this field remained unclear. This paper offered a systematic review of SML use in urban studies based on 58 publications. It recorded the topic, use case, application, data type, and method of each paper, providing statistical insights into trends and evolution. The studies were categorized into three scales, six application types, and six data types. We found that mesoscale studies are the most popular, accuracy comparison is the most common application type, and multiple data sources are the most common input. The Random Forest (RF) algorithm was the most frequently used across categories. A detailed review showed how SML is applied in various outdoor thermal comfort (OTC) studies, suggesting that future research should cover wider study areas and longer time spans. Qualitative methods should also be incorporated into SML research as complementary tools. We discussed SML methods, highlighting algorithm synergy and portability as key areas for future OTC research. The limitations of current applications, such as high computing costs, parameter adjustment complexity, and tedious pretreatment, were identified, along with development directions to address these issues. For input data, we recommended using more dimensional datasets and advanced preprocessing models to enhance prediction accuracy and depth. In conclusion, this review demonstrated SML’s potential to effectively solve outdoor thermal comfort problems. It refines and summarizes existing workflows and applications, offering three propositions for future development in this field.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".