An Ensemble Kalman Filter for Numerical Weather Prediction Based on Variational Data Assimilation: VarEnKF
Bibliographic record
Abstract
Abstract Several NWP centers currently employ a variational data assimilation approach for initializing deterministic forecasts and a separate ensemble Kalman filter (EnKF) system both for initializing ensemble forecasts and for providing ensemble background error covariances for the deterministic system. This study describes a new approach for performing the data assimilation step within a perturbed-observation EnKF. In this approach, called VarEnKF, the analysis increment is computed with a variational data assimilation approach both for the ensemble mean and for all of the ensemble perturbations. To obtain a computationally efficient algorithm, a much simpler configuration is used for the ensemble perturbations, whereas the configuration used for the ensemble mean is similar to that used for the deterministic system. Numerous practical benefits may be realized by using a variational approach for both deterministic and ensemble prediction, including improved efficiency for the development and maintenance of the computer code. Also, the use of essentially the same data assimilation algorithm would likely reduce the amount of numerical experimentation required when making system changes, since their impacts in the two systems would be very similar. The variational approach enables the use of hybrid background error covariances and may also allow the assimilation of a larger volume of observations. Preliminary tests with the Canadian global 256-member system produced significantly improved ensemble forecasts with VarEnKF as compared with the current EnKF and at a comparable computational cost. These improvements resulted entirely from changes to the ensemble mean analysis increment calculation. Moreover, because each ensemble perturbation is updated independently, VarEnKF scales perfectly up to a very large number of processors.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.012 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".