Data assimilation for weather forecasting: Are we optimizing the appropriate initial conditions?
Bibliographic record
Abstract
Weather forecasting being an initial value problem, we generally seek to minimize the error on the initial state of the model in hopes of reducing errors at future forecast times. However, if our primary goal were to reduce forecast errors, would we minimize initial condition errors differently? Forecast errors are determined by two terms: the initial condition error, and the error growth during the forecast. The first term dominates the errors for short forecast times, but as forecast time increases, the second term becomes increasingly important. Thus, we wondered if there is a way to also lower forecast error growth by minimizing initial conditions errors differently.In that context, we considered whether other possible strategies such as minimizing errors on tendencies and minimizing errors on larger more predictable scales could lead to better forecasts. We chose to answer these questions by analyzing four months of ensemble forecasts from the Canadian Global Ensemble Prediction System. Each member was successively taken to be the truth, and different forms of initial condition “errors” (in values, in tendencies, at large scales) of other members were computed and compared with forecast errors as a function of forecast time.We found that for short-range forecasts, forecast errors are better correlated with errors on initial state variables, whereas for long-range forecasts, forecast errors are better correlated with a combination of errors on the initial state variables and their tendencies. Furthermore, we found that lowering initial tendencies errors by 10% leads to better forecast improvement than lowering initial state variable errors by 10% for all forecast times. Lastly, focusing strictly on initial condition errors at larger scales did not provide better forecasts. These results suggest that having data assimilation systems also minimize errors on initial tendencies could yield better forecasts
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.012 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.002 | 0.005 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".