The Canadian Land Data Assimilation System (CaLDAS): Description and Synthetic Evaluation Study
Bibliographic record
Abstract
Abstract The Canadian Land Data Assimilation System (CaLDAS) has been developed at the Meteorological Research Division of Environment Canada (EC) to better represent the land surface initial states in environmental prediction and assimilation systems. CaLDAS is built around an external land surface modeling system and uses the ensemble Kalman filter (EnKF) methodology. A unique feature of CaLDAS is the use of improved precipitation forcing through the assimilation of precipitation observations. An ensemble of precipitation analyses is generated by combining numerical weather prediction (NWP) model precipitation forecasts with precipitation observations. Spatial phasing errors to the NWP first-guess precipitation forecasts are more effective than perturbations to the precipitation observations in decreasing (increasing) the exceedance ratio (uncertainty ratio) scores and generating flatter, more reliable ranked histograms. CaLDAS has been configured to assimilate L-band microwave brightness temperature TB by coupling the land surface model with a microwave radiative transfer model. A continental-scale synthetic experiment assimilating passive L-band TBs for an entire warm season is performed over North America. Ensemble metric scores are used to quantify the impact of different atmospheric forcing uncertainties on soil moisture and TB ensemble spread. The use of an ensemble of precipitation analyses, generated by assimilating precipitation observations, as forcing combined with the assimilation of L-band TBs gave rise to the largest improvements in superficial soil moisture scores and to a more rapid reduction of the root-zone soil moisture errors. Innovation diagnostics show that the EnKF is able to maintain a sufficient forecast error spread through time, while soil moisture estimation error improvements with increasing ensemble size were limited.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.003 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.003 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".