How accurate are individual forecasters? : an assessment of the Survey of Professional Forecasters
Bibliographic record
Abstract
This master thesis addresses the forecast accuracy of individual inflation forecasts from the Survey of Professional Forecasters. Based on a variety of accuracy statistics, there are five main findings to report of. First, I find that some individuals are able to accurately predict inflation over time, and that forecasters on average have improved their accuracy over time. Second, forecasting accuracy becomes worse during recessions compared to the average accuracy in the respective decades but accuracy have improved in newer recessions compared to old ones. Nonetheless, some individuals are able to outperform the mean and a random walk model. Third, I find no difference in accuracy among industries, but I find evidence for biased forecasts for the three and four quarter horizon. Fourth, I find evidence for bias in roughly one-third of the individuals for all forecasting horizons. These results improve slightly when only data from the last two decades are being analysed. Fifth, the majority of individuals perform significantly worse than a random walk model regardless of used time span.\nI also find several problems with the database. These includes: missing values for the one-yearahead forecast, irregularities in forecasters’ response, reallocation of used ID’s, changing base year and inconsistencies in individuals’ forecasts.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.015 | 0.073 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.003 | 0.003 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.002 | 0.003 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".