Assessing the skill of the Pacific Decadal Oscillation (PDO) in a decadal prediction experiment
Bibliographic record
Abstract
A modified approach to the assessment of the prediction skill of “modes of variability” is proposed and applied to a decadal prediction experiment. In particular, the skill of predicting the Pacific Decadal Oscillation (PDO) is investigated. The approach depends on separately calculating the EOFs of the observations, the ensemble of forecasts, and an ensemble of simulations made with the same model and external forcing. The skill of predicting and simulating the spatial structure of the modes is captured by comparing forecast and simulated EOFs with the observation-based EOFs. This is in contrast to the case where forecasts and simulations are expanded in observation-based EOFs, or other structure functions, which gives no direct information about the model-based EOF structures themselves. The skill of predicting the temporal evolution of EOFs is separately captured by comparing the associated expansion functions. Finally, the contribution of the modes to the overall prediction skill is obtained by weighting the spatial and temporal skills with the variances involved. The behaviour of the first mode, identified as the PDO, is given particular attention. Perhaps not unexpectedly, the EOF structure of the forecasts more closely resembles that of the simulations than that of the observations, but both reproduce the structure of the observed PDO quite well with spatial correlations near 0.8. The temporal correlation of the expansion functions is near 0.7 for year 1 forecasts and declines toward zero subsequently. The overall correlation skill for the North Pacific is dominated by the PDO with a small contribution from the second mode and none from the third mode.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".