Decoupling impacts of weather conditions on interannual variations in concentrations of criteria air pollutants in south China – constraining analysis uncertainties by using multiple analysis tools
Bibliographic record
Abstract
Abstract. In this study, three methods including the random forest (RF) algorithm, boosted regression trees (BRTs) and the improved complete ensemble empirical mode decomposition with adaptive noise (ICEEMDAN) were adopted for investigating emission-driven interannual variations in concentrations of air pollutants including PM2.5, PM10, O3, NO2, CO, SO2 and (NO2+O3) monitored in six cities in south China from May 2014 to April 2021. The first two methods were used to calculate the deweathered hourly concentrations, and the third one was used to calculate decomposed hourly residuals. To constrain the uncertainties in the calculated deweathered or decomposed hourly values, a self-developed method was applied to calculate the range of the deweathered percentage changes (DePCs) of air pollutant concentrations in annual scale. Emission-driven trends and emission-driven percentage changes (PCs) during the whole seven-year period were generated with the four methods being applied to analyzing the data. The consistency in the trends between the RF-deweathered and BRTs-deweathered concentrations and the ICEEMDAN-decomposed residuals of an air pollutant in a city reaches approximately 70 % of all the studied cases, but that in the PCs reaches only approximately 30 % of all the cases. The remaining cases with inconsistent trends and/or PCs indicated large uncertainties produced by one or more of the three methods. The calculated PCs from the deweathered concentrations and decomposed residuals were thus combined with the corresponding range of DePCs calculated from the self-developed method to gain the robust range of DePCs where applicable. Building on the robust ranges, the mitigation effects were discussed.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.002 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".