The Performance of Shrinkage Estimator for Stock Portfolio Selection in Case of High Dimensionality
Bibliographic record
Abstract
Harry Markowitz introduced the Modern Portfolio Theory (MPT) for the first time in 1952 which has been applied widely for optimal portfolio selection until now. However, the theory still has some limitations that come from the instability of covariance matrix input. This leads the selected portfolio from MPT model to change the status continuously and to suffer the high cost of transaction. The traditional estimator of the covariance matrix has not solved this limitation yet, especially when the dimensionality of the portfolio soars. Therefore, in this paper, we conduct a practical discussion on the feasible application of the shrinkage estimator of the covariance matrix, which is expected to encourage the investors focusing on the shrinkage–based framework for their portfolio selection. The empirical study on the Vietnam stock market in the period of 2011–2021 shows that the shrinkage approach has much better performance than other traditional methods on the primary portfolio evaluation criteria such as return, level of risk, Sharpe ratio, maximum loss, and Alpla coefficient, especially the superiority is even more evident when the dimension of covariance matrix increases. The shrinkage approach tends to create more stable and secure portfolios than other estimators, as demonstrated by the average volatility and maximum loss criteria with the lowest values. Meanwhile, the factor model approach is able to generate portfolios with higher average returns and lower portfolio turnover; and the traditional approach gives good results in the case of low—dimensionality. Besides, the shrinkage method also shows effectiveness when beating the tough market benchmarks such as VN-Index and 1/N portfolio strategy on almost performance metrics in all scenarios.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.009 | 0.033 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.002 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".