On Monthly Mean Surface Wind Speed Data Homogenization and Trend Assessment
Bibliographic record
Abstract
Using hourly surface wind speed data from 155 stations in Canada, this study first developed a homogenized monthly mean wind speed dataset for the period of 1953-2023, which was then used to characterize observed changes in surface wind speed in Canada. The hourly data were first quality controlled and adjusted for non-standard anemometer heights before being used to calculate monthly mean wind speed series. To identify artificial discontinuities, the monthly mean wind speed series were subject to a semi-automated comprehensive data homogenization procedure, which uses a combination of station metadata and multiple statistical tests with and without using reference series. Reference series used include up to four best significantly-correlated neighbour stations’ data series, the ensemble mean series of monthly wind speed taken from the Twentieth Century Reanalysis version 3 (20CRv3), and monthly mean geostraphic wind speeds derived from homogenized surface pressure data. The results from the automated procedure were then reviewed manually using metadata and visual inspection of the multiphase regression fits with expert judgement. As a result, all the 155 data series were identified to have one or more artificial discontinuities, which were diminished by quantile matching adjustments. Anemometer height change, station joining, relocation, instrument changes/problems were found to be the main causes of data inhomogeneities. The homogenized dataset for 1953-2023 shows wind stilling in region from northern British Columbia (BC) to southern Yukon-Northwest Territories and from southern Prairies to Quebec-Labrador, which was matched with wind strengthening in the region from southern-central BC to the Rocky Mountains, and in Newfoundland and the high Arctics. The trend pattern of in-situ wind speed data bear substantial similarity to that of both the  ERA5 and 20CRv3 reanalysis wind speed data.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".