Prediction of the quasi‐biennial oscillation with a multi‐model ensemble of <scp>QBO</scp>‐resolving models
Bibliographic record
Abstract
Abstract A multi‐model study is carried out to investigate the ability of models to predict the evolution of the quasi‐biennial oscillation (QBO) up to 12 months in advance. All models are initialised from common reanalysis data, and forecasts run for a common set of 30 start dates over 15 years. All models have high skill in predicting the phase evolution of the QBO at 20–30 hPa, with slightly more variable results at higher and lower levels. Other aspects of the predicted QBO are of variable quality, and in some cases are consistently poor. QBO easterlies are too weak in all models at 20–50 hPa, while westerlies can be either too strong or too weak. This results in both a reduced amplitude of the QBO and a westerly bias in zonal‐mean winds, notably at 30 hPa. At 70 hPa models tend to have reduced QBO amplitude and an easterly bias. Despite these failings, a multi‐model ensemble of bias‐ and variance‐corrected forecasts can be used to give accurate and reliable QBO forecasts up to at least a year ahead. Analysis of the zonal momentum budget during the first month of the forecast shows that large‐scale forcing from Eliassen–Palm flux divergence and vertical advection are handled fairly well by the models, although vertical advection terms tend to be weaker than reanalysis estimates. Total tendencies show common errors, suggesting common failings in gravity‐wave drag treatments. Teleconnections from the QBO to Northern Hemisphere winter circulation are also examined, and do not appear to be realistic beyond the first month. Analysis of initialised forecasts is a powerful tool for diagnosing the accuracy of model processes driving the QBO.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".