Sampling frequency of climate data for the determination of daily temperature and daily temperature extrema
Bibliographic record
Abstract
Abstract The sampling frequency of temperature data is examined. The range of sampling from hourly to twice daily are explored to determine the uncertainty that is introduced by sampling less frequently than hourly in the determination of daily temperature and daily temperature extrema. The standards for comparison for the daily average temperature are the average of hourly data and the average of the daily maximum and minimum temperature. Hourly temperature data from 12 Canadian climate stations are examined for several decades leading up to 2017. Daily average temperatures were calculated using data sampled 24 times a day (hourly), 12 times, 8 times, 6 times, 4 times and twice daily. Two triad algorithms from the literature and an experimental one are assessed relative to these sampling frequencies. The sampling frequency analysis was remarkably consistent across all climate stations. The departure from the hourly estimate ranged from 0.1°C for the bi‐hourly sampling to ~1°C for the twice daily sampling. The uncertainty associated with the min/max method consistently fell within that of three and four samples per day. Comparison of triad algorithms, based on a quantitative criterion for determination of best sampling hours, revealed a station specific triad that outperforms algorithms from the literature and thrice daily evenly spaced sampling. Minimum and maximum estimates were compared across the different sampling frequencies for all stations as well. The accuracy of estimating temperature extrema decreases with lower sampling rates with the exception of the 8 hr sampling where hour of sampling influences accuracy. The results demonstrate that the local climate characteristics needs to be considered when choosing the optimal sampling frequency and calculation method for daily means and extrema.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.014 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.002 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".