Bibliographic record
Abstract
I investigate the question of how to construct a benchmark replicating portfolio consisting of a subset of the benchmark’s components. I consider two approaches: a sequential stepwise regression and another method based on factor models of security returns’ first and second moments. The first approach produces the standard hedge portfolio that has the maximum feasible correlation with the benchmark. The second approach produces weights that are proportional to a “signal-to-noise” ratio of factor beta to idiosyncratic volatility. Using a factor model of securities returns allows the use of a larger number of securities than the number of time periods used to estimate the parameters of the factor model. I also consider a second objective that maximizes expected returns subject to a target tracking error variance. The security selection criterion naturally extends to the product of the information ratio and the signal-to-noise ratio. The optimal tracking portfolio is either a one-fund or a two-fund portfolio rule consisting of the optimal hedging portfolio, the tangent portfolio or the global minimum variance portfolio, depending on what constraints are imposed on the objective function. I construct buy-and-hold replicating portfolios using the algorithms presented in the paper to track a widely followed stock index with very good results both in-sample and out-of-sample.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.016 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.002 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.003 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".