Haar wavelets, fluctuations and structure functions: convenient choices for geophysics
Bibliographic record
Abstract
Abstract. Geophysical processes are typically variable over huge ranges of space-time scales. This has lead to the development of many techniques for decomposing series and fields into fluctuations Δv at well-defined scales. Classically, one defines fluctuations as differences: (Δvdiff = v(x+Δx)-v(x) and this is adequate for many applications (Δx is the "lag"). However, if over a range one has scaling Δv ∝ ΔxH, these difference fluctuations are only adequate when 0 < H < 1. Hence, there is the need for other types of fluctuations. In particular, atmospheric processes in the "macroweather" range ≈10 days to 10–30 yr generally have −1 < H < 0, so that a definition valid over the range −1 < H < 1 would be very useful for atmospheric applications. A general framework for defining fluctuations is wavelets. However, the generality of wavelets often leads to fairly arbitrary choices of "mother wavelet" and the resulting wavelet coefficients may be difficult to interpret. In this paper we argue that a good choice is provided by the (historically) first wavelet, the Haar wavelet (Haar, 1910), which is easy to interpret and – if needed – to generalize, yet has rarely been used in geophysics. It is also easy to implement numerically: the Haar fluctuation (ΔvHaar at lag Δx is simply equal to the difference of the mean from x to x+ Δx/2 and from x+Δx/2 to x+Δx. Indeed, we shall see that the interest of the Haar wavelet is this relation to the integrated process rather than its wavelet nature per se. Using numerical multifractal simulations, we show that it is quite accurate, and we compare and contrast it with another similar technique, detrended fluctuation analysis. We find that, for estimating scaling exponents, the two methods are very similar, yet Haar-based methods have the advantage of being numerically faster, theoretically simpler and physically easier to interpret.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.013 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.003 | 0.005 |
| Science and technology studies | 0.001 | 0.003 |
| Scholarly communication | 0.005 | 0.006 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.002 | 0.005 |
| Insufficient payload (model declined to judge) | 0.004 | 0.004 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".