Representativity of cloud‐profiling radar observations for data assimilation in numerical weather prediction
Bibliographic record
Abstract
Abstract An algorithm is presented that aims to enhance the use of satellite‐based cloud‐profiling radar (CPR) data for the purpose of NWP data assimilation. It resembles the EarthCARE mission's scene construction algorithm: off‐nadir passive radiances get spectrally matched with nadir passive radiances, and the latter's collocated CPR profile gets replicated at the former's location. This process gets repeated until all passive pixels that cover an NWP domain D have a proxy column of CPR data. Only domain‐averaged profiles of CPR reflectivity ⟨ZdB⟩ and cloud fraction Ac are sought after for NWP domains measuring, nominally, 25 × 25 km. If CPR measurements intersect D, then in addition to using just the intersecting values to represent ⟨ZdB⟩ and Ac, the full array of proxy values get used. This is referred to as local estimation. If, however, CPR measurements do not intersect D, but are not too distant, this is referred to as non‐local estimation; estimates of ⟨ZdB⟩ and Ac rest entirely on proxies. Current and planned satellite‐based CPRs have nadir‐pointing narrow fields‐of‐view, so it is difficult to see how to verify the algorithm with anything other than synthetic observations. Thus, it was assessed here with simulated cloudy atmospheres and 2D and 3D distributions of associated passive radiances and CPR reflectivities. In general terms, for local domains the algorithm performs slightly better than simply averaging intersecting CPR profiles. This at least demonstrates that the proxies are not detrimental. For non‐local domains, where there are no intersecting measurements to average, the algorithm appears to perform well enough to include two or three 25 × 25 km NWP domains either side of sequences of local domains.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.003 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".