On Loss Functions and Ranking Forecasting Performances of Multivariate Volatility Models
Bibliographic record
Abstract
A large number of parameterizations have been proposed to model conditional variance dynamics in a multivariate framework. However, little is known about the ranking of multivariate volatility models in terms of their forecasting ability. The ranking of multivariate volatility models is inherently problematic because it requires the use of a proxy for the unobservable volatility matrix and this substitution may severely affect the ranking. We address this issue by investigating the properties of the ranking with respect to alternative statistical loss functions used to evaluate model performances. We provide conditions on the functional form of the loss function that ensure the proxy-based ranking to be consistent for the true one - i.e., the ranking that would be obtained if the true variance matrix was observable. We identify a large set of loss functions that yield a consistent ranking. In a simulation study, we sample data from a continuous time multivariate diffusion process and compare the ordering delivered by both consistent and inconsistent loss functions. We further discuss the sensitivity of the ranking to the quality of the proxy and the degree of similarity between models. An application to three foreign exchange rates, where we compare the forecasting performance of 16 multivariate GARCH specifications, is provided. Un grand nombre de méthodes de paramétrage ont été proposées dans le but de modéliser la dynamique de la variance conditionnelle dans un cadre multivarié. Toutefois, on connaît peu de choses sur le classement des modèles de volatilité multivariés, du point de vue de leur capacité à permettre de faire des prédictions. Le classement des modèles de volatilité multivariés est forcément problématique du fait qu'il requiert l'utilisation d'une valeur substitutive pour la matrice de la volatilité non observable et cette substitution peut influencer sérieusement le classement. Nous abordons ce problème en examinant les propriétés du classement en relation avec les fonctions de perte statistiques alternatives utilisées pour évaluer la performance des modèles. Nous présentons des conditions liées à la forme fonctionnelle de la fonction de perte qui garantissent que le classement fondé sur une valeur de substitution est constant par rapport au classement réel, c'est-à-dire à celui qui serait obtenu si la matrice de variance réelle était observable. Nous établissons un vaste ensemble de fonctions de perte qui produisent un classement constant. Dans le cadre d'une étude par simulation, nous fournissons un échantillon de données à partir d'un processus de diffusion multivarié en temps continu et comparons l'ordre généré par les fonctions de perte constantes et inconstantes. Nous approfondissons la question de la sensibilité du classement à la qualité de la substitution et le degré de ressemblance entre les modèles. Une application à trois taux de change est proposée et, dans ce contexte, nous comparons l'efficacité de prédiction de 16 paramètres du modèle GARCH multivarié (approche d'hétéroscédasticité conditionnelle autorégressive généralisée).
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.030 | 0.070 |
| Meta-epidemiology (narrow) | 0.002 | 0.000 |
| Meta-epidemiology (broad) | 0.002 | 0.001 |
| Bibliometrics | 0.003 | 0.002 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.003 | 0.004 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.002 | 0.003 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".