Distribution-Oriented Verification of Limited-Area Model Forecasts in a Perfect-Model Framework
Bibliographic record
Abstract
Nested limited-area models (LAMs) have been used by the scientific community for a long time, with the implicit assumption that they are able to generate meaningful small-scale features that were absent in the lateral boundary conditions and sometimes even in the initial conditions. This hypothesis has never been seriously challenged in spite of reservations expressed by part of the scientific community. In order to study this hypothesis, a perfect-model approach is followed. A high-resolution LAM driven by global analyses is used over a large domain to generate a “reference run.” These fields are filtered afterward to remove small scales in order to mimic low-resolution nesting data. The same high-resolution LAM, but over a small domain, is nested with these filtered fields and run for several days. The ability of the LAM to regenerate the small scales that were absent in the initial and lateral boundary conditions is estimated by comparing both runs over the same region. The simulations are analyzed for several variables using a distribution-oriented approach, which provides an estimation of the forecasting ability as a function of the value of the variable. It is found that variables with steep spectra, such as geopotential and temperature, display good forecasting skills for the entire range of values but improve little the forecast skill of a low-resolution perfect model. For noisier variables with flatter spectra, such as vorticity and precipitation, the high-resolution forecast provides a more realistic and extended range of forecast values for the variables, but rather low skill for extreme events. The probability of a successful forecast for these extreme cases, however, is much higher than that of a random model. When errors in the phase in the weather systems are not penalized, forecasting skill increases considerably. This suggests that, despite the inability to perform as pointwise deterministic forecasts, useful information may be generated by LAMs if considered in a probabilistic way.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".