A Hybrid Statistical Downscaling Framework Based on Nonstationary Time Series Decomposition and Machine Learning
Bibliographic record
Abstract
Abstract Downscaling techniques are effective to bridge the scale gap between global circulation models and regional studies. Statistical downscaling methods are prevalent due to their advantages in high computational efficiency and accuracy. However, an implicit assumption of most statistical techniques is that time series should be relatively stationary after certain function transformations. Otherwise, statistics of nonstationary time series may be meaningless for describing future behavior. In this study, a hybrid statistical downscaling framework was developed through integrating bivariate empirical mode decomposition (BEMD) and a machine learning method to extract multi‐timescale features from nonstationary data to enhance downscaling performance. The proposed framework can reduce the effects of non‐stationarity in data‐driven models by using the BEMD method, which can decompose time series into independent and stationary components at multiple time‐frequency resolutions. It was applied to downscale monthly precipitation and temperature of Canadian Earth System Model in multiple stations with different climate types in the Central Valley of California, USA, to verify its accuracy and generalization ability. The performance of downscaling maximum and minimum temperatures (R2 > 0.9) was more accurate than that of precipitation. The potential reason is that precipitation is more sensitive to transient weather phenomena, which can only be extracted from data with higher temporal resolution. The proposed model was further compared with models based on discrete wavelet transform and models without time series decomposition. The results showed that a decomposition strategy of the proposed framework can improve the downscaling accuracy, potentially providing a viable option to deal with the nonstationary of data in statistical downscaling models.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".