Unified scaling framework for Holocene, Quaternary and Phanerozoic geochronology variability
Bibliographic record
Abstract
With few exceptions, paleodata are irregularly sampled; this poses numerous challenges for the statistical characterization of paleoindicators, this includes the indicators needed to understand the climate and macroevolution. The key variable is the measurement density - the number of measurements per unit time (r(t)). Our study used 27 paleoindicators collectively spanning time scales from years to hundreds of millions of years.Using Haar fluctuation analysis and for all the series, we show that r(t) has two scaling regimes. At high frequencies, there is a low intermittency (quasi-Gaussian) scaling regime (intermittency parameter C1 ≈ 0). Over this regime, the fluctuation exponent H is negative implying that the chronologies become more uniform at longer time scales, r(t) is commonly close to a Gaussian white noise (H = -1/2). In contrast, at low frequencies, r(t) is highly intermittent (large C1), but it also has positive H so that fluctuations tend to grow with scale but in a highly intermittent fashion. In this this regime, “gaps” at all scales are important. The two regimes have simple physical interpretations: the high frequency behaviour can be explained by fairly smooth (but scaling) sedimentation rates, whereas the low frequencies can be explained by scaling erosion processes that introduce gaps over a wide range of scales (in conformity with the Sadler effect). To confirm this interpretation, we introduce a simple multiplicative sedimentation - erosion model that is close to the data. Finally, we empirically show that the gaps typically have extreme power law probability tails so that the series are not only scaling in time, but also in probability space.A key issue for paleontologists is the effect of variable r(t) on the paleoindicator estimates themselves (e.g. on paleotemperatures T(t)). Using Haar fluctuations we determined the fluctuation - fluctuation correlation R(Δt) = < Δ r(Δt) ΔT(Δt) >. When R(Δt) is small, the measurements and indicators are statistically independent so that the biases due to r(t) variability on paleoindicator statistics are easy to correct. However, at large Δt, the correlations are frequently large, and this poses additional difficulties in data interpretation. Strong correlations were observed in the Quaternary, but not the Holocene or Phanerozoic.Our study spans more than 8 orders of magnitude in time scale and it shows that it is wrong to theorize paleoseries as being fundamentally regularly sampled but interspersed with occasional data “holes” that can be dealt with using conventional techniques such as interpolation. While Haar fluctuation analysis is insensitive to the chronology variability - and if needed can easily be statistically corrected for any biases that it introduces - this is not true of existing spectral estimators that are extremely sensitive to scaling data gaps.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.007 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.003 | 0.002 |
| Science and technology studies | 0.001 | 0.003 |
| Scholarly communication | 0.004 | 0.004 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.006 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".