Is there an Optimal Forecast Combination? A Stochastic Dominance Approach to Forecast Combination Puzzle
Bibliographic record
Abstract
Even though different optimal forecast combination weights are offered for static, dynamic, or time-varying situations, empirical findings support the simple average forecast combination outperforms more sophisticated weighting schemes and/or the best individual model. Using an approach that relies on consistent tests for stochastic dominance efficiency, an alternative optimal weighting scheme is proposed. These tests are considered for a given forecast combina-tion (i.e. equal weighted average of forecasts) with respect to all possible forecast combinations constructed from a set of individual forecasts to obtain optimal or worst forecast combina-tions. In our empirical applications we find that equally weighted forecast combinations are neither optimal nor the worst forecast combination. For the optimal forecast combination, the best forecasting model, i.e. the model which assumes relatively more weight than other forecast models, differs with the variable being forecasted and for different forecast horizons. On the other hand, random walk is the model that consistently contributes with more than arbitrarily assigned equal weights for the worst forecast combination for all variables being forecasted and for all forecast horizons.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.002 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".