Leveraged and inverse ETF performance during the financial crisis
Bibliographic record
Abstract
Purpose Leveraged and inverse ETFs (hereafter leveraged ETFs) have received much press coverage of late due to issues with their performance. Managers and the media have focused investors' attention on the impact of compounding, when the funds are held for more than one day. The aim of this paper is to lay out a framework for assessing the performance of leveraged ETFs. Design/methodology/approach The authors propose a simple way to disentangle the effect of compounding and that of the management of the fund and the trading premiums/discounts, all of which affect investors' bottom line. The former is influenced by the effectiveness and the costs of the manager's (synthetic) replication strategy and the use of leverage. The latter reflects liquidity and the efficiency of the market. Findings The paper finds that tracking errors were not caused by the effects of compounding alone. Depending on the fund, the impact of management factors can outweigh the impact of compounding, and substantial premiums/discounts caused by reduced liquidity during the financial crisis further distorted performance. Originality/value The authors propose a framework for practitioners to evaluate the performance of leveraged ETFs. This framework highlights a very topical issue, that of the impact of synthetic replication, which all leveraged ETFs use. Financial regulators such as the SEC and the Financial Stability Board have all taken issue with synthetically replicated ETFs. In leveraged ETFs, this issue is masked by the effects of compounding. The framework the authors propose allows investors to disentangle the two effects.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.023 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.002 | 0.002 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.003 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".