Data-driven and physics-based modeling approaches and their integration in building digital twins: A systematic review
Notice bibliographique
Résumé
Interest in digital twin technology has grown significantly within the building sector as part of the broader digital transformation in the architecture, engineering, and construction industry. A building digital twin is a virtual replica that captures a building’s static and dynamic behavior through data, information, and models. Digital twin models can be developed using data-driven or physics-based approaches, each with distinct advantages and limitations. Data-driven models can learn complex behaviors from data and scale well, but they require large datasets and often lack interpretability. In contrast, physics-based models offer interpretability and generalizability through fundamental principles but can be computationally demanding. Consequently, building digital twins can benefit greatly from integrating both approaches through hybrid modeling. However, the literature lacks a comprehensive analysis of integration strategies within building digital twins. This study addresses that gap by reviewing advances in data-driven and physics-based modeling and analyzing various integration levels. The results show that most studies rely on siloed models, using either approach independently without leveraging their complementary strengths. Some adopted sequential integration, where one model informs the other but lacks real-time or iterative feedback. A few achieved coupled integration, involving active data exchange and collaboration between models. Only three studies explored fusion integration, where both approaches are fully unified into a single model. Based on this review, a method is proposed for selecting the appropriate level of integration, considering factors such as data availability, interpretability, generalizability, and domain knowledge. Finally, key research gaps and future directions are identified to guide further work. • Reviews data-driven and physics-based modeling approaches in building DTs. • Examines varying levels of integration between data-driven and physics-based models. • Discusses the key trade-offs for each modeling approach and integration level. • Presents guidelines for selecting the appropriate integration level. • Identifies key research gaps to direct future research efforts.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,012 | 0,045 |
| Méta-épidémiologie (sens strict) | 0,002 | 0,001 |
| Méta-épidémiologie (sens large) | 0,004 | 0,007 |
| Bibliométrie | 0,013 | 0,012 |
| Études des sciences et des technologies | 0,001 | 0,001 |
| Communication savante | 0,004 | 0,005 |
| Science ouverte | 0,003 | 0,002 |
| Intégrité de la recherche | 0,002 | 0,002 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,005 | 0,001 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».