Grid Search Optimized Machine Learning based Modeling of CO2 Emissions Prediction from Cars for Sustainable Environment
Bibliographic record
Abstract
Carbon emissions have increased dramatically because of industrialization, trapping heat in the atmosphere and hastening climate change. This is a serious threat to the wealth, security, and well-being of the world. The effects are extensive, ranging from severe weather, disease outbreaks, and economic disruption to food insecurity and water scarcity. The World Health Organization (WHO) has determined that climate change poses the greatest threat to public health in the twenty-first century. Thus, precise CO2 emissions have emerged as a crucial concern in recent times. Several studies have tried to forecast the amount CO2 from industry and power plant using statistical analysis. Efficiency, robustness and diverse application was the limitation of the study. In this study, we have proposed an AI based model that is able to predict the amounts of CO2 emissions from cars. We applied a grid search-optimized machine learning approach using the publicly available Canadian dataset. Incorporation of different statistical analyses and preprocessing techniques such as duplicate data management, outlier rejection, scaling contributed to enhance the quality of the dataset. Later, grid search techniques were applied to tune the KNN, RF, and SVR models. The approach has enhanced the performance of CO2 emissions prediction. In the study, we further used the explainability of the random forest model to check the bias and fairness of predictability. MSE, RMSE, and R-squared metrics of the proposed approach were the highest as the state of the art.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".