LSTM and Transformer-based framework for bias correction of ERA5 hourly wind speeds
Bibliographic record
Abstract
Reanalysis-derived wind speeds are central to large-scale wind resource assessment (WRA). However, their coarse spatial resolution often introduces significant biases, particularly in complex terrains and coastal areas. A deep learning (DL) framework using LSTMs or Transformers was introduced to correct systematic biases and temporal variability in reanalysis-derived wind speeds by modeling a time-resolved scaling factor, which is then used to adjust ERA5 wind speeds. The proposed framework's spatiotemporal generalization capability was rigorously evaluated using a test set of 170 independent stations distributed throughout Canada in diverse environmental conditions. Results showed that the DL framework outperformed a standard bias correction method based on the Global Wind Atlas. It improved the median wind speed, the temporal variability, and the probability distributions of ERA5 wind speeds in coastal areas and complex terrains. Specifically, in coastal regions, the DL models increased the explained variability of median wind speed by over 70% relative to ERA5. In regions characterized by high surface roughness length, such as forests and urban areas, these models achieved average improvements of more than 10% in MAE and RMSE of the time series. While the DL models performed well in representing the probability distribution of the most typical wind speed values, some challenges remain in improving the distribution of extreme wind speeds. Overall, this framework represents a promising advancement in enhancing the accuracy of reanalysis-derived wind speeds in large-scale WRA. By reducing biases in ERA5 wind data, the DL framework supports more reliable site selection and estimation of long-term energy production.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".