A Hybrid Deep Learning Model for Forecasting PM2.5 Concentrations in Northern Thailand from Satellite Images
Bibliographic record
Abstract
Air pollution is a significant environmental issue with extensive impacts, particularly concerning particulate matter smaller than 2.5 microns (PM2.5), which poses serious public health risks,especially respiratory diseases such as various diseases, ischemic heart disease, strokes, chronic obstructive pulmonary disease, tracheal, bronchus, lung cancer, and even increased premature death rates. Northern Thailand is one of the areas with the most severe PM2.5 problems, especially during the summer (February to May), primarily due to the large amount of agricultural field burning and forest fires by ethnic groups after the harvest season.This research proposes a hybrid model of Convolution Neural Network (CNN) and Long Short-Term Memory (LSTM) for PM2.5 concentration forecasting using satellite images of four environmental variables: aerosol optical depth, temperature, precipitation, and ozone. These variables are important factors in the occurrence of PM2.5. The efficiency of the CNN-LSTM model was assessed by comparing performance with classification deep learning models (CNN, LSTM), Seasonal Autoregressive Integrated Moving Average with Exogenous Variables (SARIMAX), and Multiple Linear Regression (MLR). The findings indicate that The CNN-LSTM model achieves higher accuracy than the other models, achieving an R2 of 98.38%, MAPE of 2.47%, and significantly lower RMSE (3.0672 μg/m3) and MAE (0.8560 μg/m3). In conclusion, this research highlights the important implications of supporting government policy formulation and public preparedness to address the PM2.5 problem, which varies in severity across seasons.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".