Ensemble Size, Balance, and Model-Error Representation in an Ensemble Kalman Filter*
Bibliographic record
Abstract
The ensemble Kalman filter (EnKF) has been proposed for operational atmospheric data assimilation.Some outstanding issues relate to the required ensemble size, the impact of localization methods on balance, and the representation of model error.To investigate these issues, a sequential EnKF has been used to assimilate simulated radiosonde, satellite thickness, and aircraft reports into a dry, global, primitive-equation model.The model uses the simple forcing and dissipation proposed by Held and Suarez.It has 21 levels in the vertical, includes topography, and uses a 144 ϫ 72 horizontal grid.In total, about 80 000 observations are assimilated per day.It is found that the use of severe localization in the EnKF causes substantial imbalance in the analyses.As the distance of imposed zero correlation increases to about 3000 km, the amount of imbalance becomes acceptably small.A series of 14-day data assimilation cycles are performed with different configurations of the EnKF.Included is an experiment in which the model is assumed to be perfect and experiments in which model error is simulated by the addition of an ensemble of approximately balanced model perturbations with a specified statistical structure.The results indicate that the EnKF, with 64 ensemble members, performs well in the present context.The growth rate of small perturbations in the model is examined and found to be slow compared with the corresponding growth rate in an operational forecast model.This is partly due to a lack of horizontal resolution and partly due to a lack of realistic parameterizations.The growth rates in both models are found to be smaller than the growth rate of differences between forecasts with the operational model and verifying analyses.It is concluded that model-error simulation would be important, if either of these models were to be used with the EnKF for the assimilation of real observations.*This paper is dedicated to the memory of Dr. Roger Daley, whose scientific insight and congeniality will be sorely missed.The understanding of balance in atmospheric models and the application of Kalman filter theory to atmospheric data assimilation are two of the areas in which Dr. Daley made important contributions.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.009 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.006 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".