Empirical study of value‐at‐risk and expected shortfall models with heavy tails
Bibliographic record
Abstract
Purpose This paper aims to test empirically the performance of different models in measuring VaR and ES in the presence of heavy tails in returns using historical data. Design/methodology/approach Daily returns of popular indices (S&P500, DAX, CAC, Nikkei, TSE, and FTSE) and currencies (US dollar vs Euro, Yen, Pound, and Canadian dollar) for over ten years are modeled with empirical (or historical), Gaussian, Generalized Pareto (peak over threshold (POT) technique of extreme value theory (EVT)) and Stable Paretian distribution (both symmetric and non‐symmetric). Experimentation on different factors that affect modeling, e.g. rolling window size and confidence level, has been conducted. Findings In estimating VaR, the results show that models that capture rare events can predict risk more accurately than non‐fat‐tailed models. For ES estimation, the historical model (as expected) and POT method are proved to give more accurate estimations. Gaussian model underestimates ES, while Stable Paretian framework overestimates ES. Practical implications Research findings are useful to investors and the way they perceive market risk, risk managers and the way they measure risk and calibrate their models, e.g. shortcomings of VaR, and regulators in central banks. Originality/value A comparative, thorough empirical study on a number of financial time series (currencies, indices) that aims to reveal the pros and cons of Gaussian versus fat‐tailed models and Stable Paretian versus EVT, in estimating two popular risk measures (VaR and ES), in the presence of extreme events. The effects of model assumptions on different parameters have also been studied in the paper.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.010 | 0.040 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.002 | 0.003 |
| Open science | 0.002 | 0.001 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".