Partial validation of a lossy compression approach to computing radiative transfer in cloud system‐resolving models
Bibliographic record
Abstract
Abstract Cloud system‐resolving models (CSRMs) routinely calculate radiative flux profiles via the Independent Column Approximation (ICA). The ICA applies 1D radiative transfer models (RTMs) to all columns in a CSRM's domain. For this study, the Partitioned Gauss–Legendre Quadrature (PGLQ) method replaced the ICA in a CSRM. The PGLQ applies RTMs to columns, identified with GLQ rules, and distributes their flux profiles to the other columns. The PGLQ approach is likened to a lossy compression algorithm that trades information for efficiency. While verification and validation of an audio compression algorithm rest, respectively, on file size reduction and sound quality according to listeners, for the PGLQ they rest on increasing while maintaining, according to experimenters, the integrity of CSRM simulations. A CSRM was run for 80 days in radiative‐convective equilibrium (1,024 × 1,024 columns and horizontal grid‐spacing of 0.25 km) for sea‐surface temperature SST = 295 and 300 K; the last 40 days were analysed. Simulations using the ICA represent the control; experiments used the PGLQ calling the RTMs fRT = 200, 5,000 and 50,000 fewer times than the control. For fRT = 50,000, several key variables spanning time/domain‐averaged cloud and radiation properties, a measure of cloud (convection) aggregation, and horizontal fluctuations of cloud and radiation fields, differ significantly from the control. In contrast, corresponding differences between the control and PGLQ with fRT ≤ 5,000, for both a given SST and differences between SSTs, are often minor, in all respects, and appear to be drawn from a single population. While these results partially validate the PGLQ method for fRT ≤ 5,000, they also indicate that overly large reductions in radiative information are detrimental.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.014 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.002 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".