A cross-checked global monthly weather station database for precipitation covering the period 1901 to 2010
Bibliographic record
Abstract
This database entry represents a comprehensive compilation of monthly weather station records for precipitation from multiple data sources for the period 1901-2010, with an emphasis on climate normal averages for the period 1961-1990. The database corresponds to the journal publication: Castellanos-Acuña, D. and Hamann, A. 2020. A cross-checked global monthly weather station database for precipitation covering the period 1901 to 2010. Geoscience Data Journal (https://rmets.onlinelibrary.wiley.com/journal/20496060, article in press, January 2020). We use digital elevation models and nearby stations to search for inconsistencies in reported station locations and recorded precipitation values. We also estimated missing values in weather station time series using a linear model approach based on interpolated anomaly surfaces. The resulting station records were ranked into ten classes, according to the completeness of records, the reliability of missing value estimations and other criteria. We corrected incomplete or erroneous location and elevation information for 12% of all available station records. A total of 23% of monthly records that had missing values could be estimated with high or moderate confidence. We sub-sampled our global database of more than 80,000 stations with various spatial filters, so that only the highest quality station for a given area was retained. Our contribution significantly enhances global data coverage compared to individual databases currently available. Even when accepting only the stations within the top two quality ranks in our combined database, and applying the coarsest spatial filter of one station per approximately 1,600 km², the remaining station count of more than 20,000 stations exceeds the largest alternative database (without a spatial filter applied) by more than 50%. The database contains a "Station Statistics" file with various flags indicating station quality and completeness of records. Monthly precipitation data is provided as one large file, but also broken down into regional files with less than one million rows each. Climate normal estimates for the 1961-1990 period, useful as a baseline prior to significant anthropogenic warming, are provided in multiple files with global coverage, but with different spatial filters applied that select the highest quality stations based for a global grid at different resolutions.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.002 | 0.000 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.003 | 0.002 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".