Using Machine Learning to Generate a GISS ModelE Calibrated Physics Ensemble (CPE)
Bibliographic record
Abstract
A neural network (NN) surrogate of the NASA GISS ModelE atmosphere (version E3) is trained on a perturbed parameter ensemble (PPE) spanning 45 physics parameters and 36 outputs. The NN is leveraged in a Markov Chain Monte Carlo (MCMC) Bayesian parameter inference framework to generate a second posterior constrained ensemble coined a “calibrated physics ensemble”, or CPE. The CPE members are characterized by diverse parameter combinations and are, by definition, close to top-of-atmosphere radiative balance, and must broadly agree with numerous hydrologic, energy cycle and radiative forcing metrics simultaneously. Global observations of numerous cloud, environment, and radiation properties (provided by global satellite products) are crucial for CPE generation. The inference framework explicitly accounts for discrepancies (or biases) in satellite products during CPE generation. We demonstrate that product discrepancies strongly impact calibration of important model parameter settings (e.g., convective plume entrainment rates; fall speed for cloud ice). Structural improvements new to E3 are retained across CPE members (e.g., stratocumulus simulation). Notably, the framework improved the simulation of shallow cumulus and Amazon rainfall while not degrading radiation fields, an upgrade that neither default parameters nor Latin Hypercube parameter searching achieved. Analyses of the initial PPE suggested several parameters were unimportant for output variation. However, many “unimportant” parameters were needed for CPE generation, a result that brings to the forefront how parameter importance should be determined in PPEs. From the CPE, two diverse 45-dimensional parameter configurations are retained to generate radiatively-balanced, auto-tuned atmospheres that were used in two E3 submissions to CMIP6
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.006 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".