Arctic Holocene proxy climate database – new approaches to assessing geochronological accuracy and encoding climate variables
Bibliographic record
Abstract
Abstract. We present a systematic compilation of previously published Holocene proxy climate records from the Arctic. We identified 170 sites from north of 58° N latitude where proxy time series extend back at least to 6 cal ka (all ages in this article are in calendar years before present – BP), are resolved at submillennial scale (at least one value every 400 ± 200 years) and have age models constrained by at least one age every 3000 years. In addition to conventional metadata for each proxy record (location, proxy type, reference), we include two novel parameters that add functionality to the database. First, "climate interpretation" is a series of fields that logically describe the specific climate variable(s) represented by the proxy record. It encodes the proxy–climate relation reported by authors of the original studies into a structured format to facilitate comparison with climate model outputs. Second, "geochronology accuracy score" (chron score) is a numerical rating that reflects the overall accuracy of 14C-based age models from lake and marine sediments. Chron scores were calculated using the original author-reported 14C ages, which are included in this database. The database contains 320 records (some sites include multiple records) from six regions covering the circumpolar Arctic: Fennoscandia is the most densely sampled region (31% of the records), whereas only five records from the Russian Arctic met the criteria for inclusion. The database contains proxy records from lake sediment (60%), marine sediment (32%), glacier ice (5%), and other sources. Most (61%) reflect temperature (mainly summer warmth) and are primarily based on pollen, chironomid, or diatom assemblages. Many (15%) reflect some aspect of hydroclimate as inferred from changes in stable isotopes, pollen and diatom assemblages, humification index in peat, and changes in equilibrium-line altitude of glaciers. This comprehensive database can be used in future studies to investigate the spatio-temporal pattern of Arctic Holocene climate changes and their causes. The Arctic Holocene data set is available from NOAA Paleoclimatology.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.013 | 0.040 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.014 | 0.018 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.006 | 0.005 |
| Open science | 0.003 | 0.003 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".