Several Mathematical Problems in Investment Management
Bibliographic record
Abstract
This thesis studies four mathematical problems in investment management. All four problems arise from practical challenges and are data-driven. \n \nChapter 2 investigates the Kelly portfolio strategy. The full Kelly strategy's deficiency in the face of estimation errors in practice can be mitigated by fractional or shrinkage Kelly strategies. This chapter provides an alternative, the RL Kelly strategy, based on a reinforcement learning (RL) framework. RL algorithms are developed for the practical implementation of the RL Kelly strategy. Extensive simulation studies are conducted, and the results confirm the superior performance of the RL Kelly strategies. \n \nIn Chapter 3, we study the discrete-time mean-variance problem under an RL framework. The continuous-time problem was theoretically studied by the existing literature but was subject to a discretization error in implementations. We compare our discrete-time model with the continuous-time model in terms of theoretical results and numerical performance. In a daily trading market setting, we find both discrete-time and continuous-time models achieve comparable performance. However, the discrete-time model outperforms better than the continuous-time model when the trading is less frequent. Our discrete-time model is not subject to the discretization error. \n \nChapter 4 explores the valuation problem of large variable annuity (VA) portfolios. A computationally appealing methodology for the valuation of large VA portfolios is a metamodelling framework that evaluates a small set of representative contracts, fits a predictive model based on these computed values, and then extrapolates the model to estimate the values of the remaining contracts. This chapter proposes a new two-phase procedure for selecting representative contracts. The representatives from the first phase are determined using contract attributes as in existing metamodelling approaches, but those in the second phase are chosen by utilizing the information contained in the values of the representatives from the first phase. Two numerical studies confirm that our two-phase selection procedure improves upon conventional approaches from the existing literature. \n \nChapter 5 focuses on the capture ratio which is a widely-used investment performance measure. We study the statistical problem of estimating the capture ratio based on a finite number of observations of a fund's returns. We derive the asymptotic distribution of the estimator, and use it for testing whether one fund has a capture ratio that is statistically significantly higher than another. We also perform hypothesis tests with real-world hedge fund data. Our analysis raises concerns regarding the models and sample sizes used for estimating capture ratios in practice.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.010 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.002 | 0.004 |
| Scholarly communication | 0.004 | 0.006 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.004 | 0.006 |
| Insufficient payload (model declined to judge) | 0.010 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".