A machine learning methodology for developing microscopic vehicular fuel consumption and emission models for local conditions using real-world measures
Notice bibliographique
Résumé
Road transport is a major contributor to world energy consumption and emissions. The validity of models developed for environmental assessment of transport projects when used out of their origins is questionable as they are only validated for the prevailing conditions at their origin. This study starts by the validation of one of the most popular transportation environmental assessment models, MOVES, for use in non-U.S. regions such as Canada through performing on-road measurements. Distinct differences between the ground-truth and MOVES predictions are revealed. MOVES underestimates fuel and CO2 rates by 17% and 35%, respectively. Nitrogen Oxides (NOx) and Particulate Matters (PM) predictions set overestimation records of up to +420%. Furthermore, MOVES output is biased for vehicle groups with specific attributes. The results of MOVES validation emphasized the need for using alternative local fuel and emission models. However, many of the existing vehicular fuel and emission modeling methodologies are criticized in aspects such as ignoring real-world training data, low diversity of test fleet, impracticality in real-world applications (such as instrument-independent eco-driving or use alongside with traffic microsimulation), and low prediction power in the non-linear multi-dimensional space of fuel consumption and emission generation. Hence, a machine learning modeling methodology relying on on-road data from a fleet of 35 vehicles is proposed. The accuracy of the proposed instrument-independent models is tried to be improved by introducing estimates of influential engine variables to the feature set through a cascaded modeling procedure. As a result, the R-squared metric reached 83%, while score improvements as high as 37% are achieved depending on the vehicle class and the machine learning technique used.Despite the considerable scores achieved by utilizing fully-connected neural networks architectures, use of techniques compatible with the serially-correlated nature of vehicular operation seems more promising in achieving higher accuracy and robustness. Moreover, generalizing the models developed for particular vehicles to more aggregate levels is a need for diversifying models’ use cases. To this end, a two-stage ensemble learning methodology based on vehicle-specific Recurrent Neural Network (RNN) models is proposed.Long Short-Term Memory (LSTM) cell architecture resulted in the best lag-specific modeling scores (compared to the other RNN cell types). Vehicle-specific ensemble models developed by combining predictions from lag-specific RNN models showed score improvement records of up to 28% compared to the best component model (4% on average). In addition, the category-specific ensembles developed on top of metamodels achieved score improvements of up to 32% compared to the best component metamodel (6% on average). Linear regression dominantly resulted in the best score improvements for NOx and PM rates at both forecast combination stages, while random forests and gradient boosting methods dominantly worked the best for fuel and CO2 rates
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,001 | 0,003 |
| Méta-épidémiologie (sens strict) | 0,001 | 0,000 |
| Méta-épidémiologie (sens large) | 0,001 | 0,001 |
| Bibliométrie | 0,001 | 0,001 |
| Études des sciences et des technologies | 0,000 | 0,000 |
| Communication savante | 0,001 | 0,001 |
| Science ouverte | 0,001 | 0,001 |
| Intégrité de la recherche | 0,001 | 0,001 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,001 | 0,001 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».