Statistical–Dynamical Seasonal Prediction Based on Principal Component Regression of GCM Ensemble Integrations
Bibliographic record
Abstract
A statistical approach to correct a dynamical ensemble forecast of future seasonal means based on the past performance of a general circulation model (GCM) is formulated. The approach combines principal component (PC) analysis with the regression technique to remove the systematic structural (distortion) error from the GCM ensemble. The performance of this statistical–dynamical method is assessed by comparing its cross-validated skill with the explicit skill achieved in the raw GCM ensembles. When the PC regression technique is applied to seasonal means from an ensemble of the Center for Ocean–Land–Atmosphere Studies (COLA) GCM, it not only recovers most of the explicit skill in the original ensemble mean, but also acts to correct significant errors in the ensemble. It is shown that some ensemble errors are due to noise that can be easily removed by applying a simple regression scheme. A novel aspect of the PC regression technique, however, is that it goes beyond the simple filtering of noise and is able to correct systematic errors in the structure of predicted fields. Thus, it has the ability to diagnose and make use of implicit skill. In the authors' application, this skill appears in the extratropical western Pacific and in east Asia, and leads to significant improvement of seasonal forecast skill. To make the PC regression scheme operationally useful, the authors develop a screening procedure for selecting skillful PCs as predictors for the regression equation. The predictors are chosen by the screening procedure based on their cross-validated performance within the training data over the whole domain or over a specified regional domain. When the procedure is applied to the COLA ensemble over the whole domain of the Northern Hemisphere, it achieves significant skill that is close to its upper bound achievable only through a postprocessing procedure. The authors also present a regional down-scaling exercise focused over eastern Canada and the northeast United States. This exercise reveals some nonlinear, asymmetric atmospheric responses to the ENSO forcing. Applications of the PC regression scheme to ensembles generated by other GCMs are also discussed. It is clear that when the SST-forced signal in the GCM is either very weak or not easily separated from noise, the regression scheme proposed will not be very successful.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.007 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".