Bibliographic record
Abstract
We explore the question of optimal aggregation level for stress testing models when the stress test is specified in terms of aggregate macroeconomic variables, but the underlying performance data are available at a loan level. Using standard model performance measures, we ask whether it is better to formulate models at a disaggregated level (“bottom up”) and then aggregate the predictions in order to obtain portfolio loss values or is it better to work directly with aggregated models (“top down”) for portfolio loss forecasts. We study this question for a large portfolio of home equity lines of credit. We conduct model comparisons of loan-level default probability models, county-level models, aggregate portfolio-level models, and hybrid approaches based on portfolio segments such as debt-to-income (DTI) ratios, loan-to-value (LTV) ratios, and FICO risk scores. For each of these aggregation levels we choose the model that fits the data best in terms of in-sample and out-of-sample performance. We then compare winning models across all approaches. We document two main results. First, all the models considered here are capable of fitting our data when given the benefit of using the whole sample period for estimation. Second, in out-of-sample exercises, loan-level models have large forecast errors and underpredict default probability. Average out-of-sample performance is best for portfolio and county-level models. However, for portfolio level, small perturbations in model specification may result in large forecast errors, while county-level models tend to be very robust. We conclude that aggregation level is an important factor to be considered in the stress-testing model design.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".