Using Machine Learning to Generate a GISS ModelE Calibrated Physics Ensemble (CPE)
Bibliographic record
Abstract
A neural network (NN) surrogate of the NASA GISS ModelE atmosphere (version E3) is trained on a perturbed parameter ensemble (PPE) spanning 45 physics parameters and 36 outputs. The NN is leveraged in a Markov Chain Monte Carlo (MCMC) Bayesian parameter inference framework to generate a second posterior constrained ensemble coined a “calibrated physics ensemble”, or CPE. The CPE members are characterized by diverse parameter combinations and are, by definition, close to top-of-atmosphere radiative balance, and must broadly agree with numerous hydrologic, energy cycle and radiative forcing metrics simultaneously. Global observations of numerous cloud, environment, and radiation properties (provided by global satellite products) are crucial for CPE generation. The inference framework explicitly accounts for discrepancies (or biases) in satellite products during CPE generation. We demonstrate that product discrepancies strongly impact calibration of important model parameter settings (e.g., convective plume entrainment rates; fall speed for cloud ice). Structural improvements new to E3 are retained across CPE members (e.g., stratocumulus simulation). Notably, the framework improved the simulation of shallow cumulus and Amazon rainfall while not degrading radiation fields, an upgrade that neither default parameters nor Latin Hypercube parameter searching achieved. Analyses of the initial PPE suggested several parameters were unimportant for output variation. However, many “unimportant” parameters were needed for CPE generation, a result that brings to the forefront how parameter importance should be determined in PPEs. From the CPE, two diverse 45-dimensional parameter configurations are retained to generate radiatively-balanced, auto-tuned atmospheres that were used in two E3 submissions to CMIP6
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.003 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".