How Many Stocks Are Sufficient for Equity Portfolio Diversification? A Review of the Literature
Bibliographic record
Abstract
Using extensive and comprehensive databases to select a subset of research papers, we aim to critically analyze previous empirical studies to identify certain patterns in determining the optimal number of stocks in well-diversified portfolios in different markets, and to compare how the optimal number of stocks has changed over different periods and how it has been affected by market turmoil such as the Global Financial Crisis (GFC) and the current COVID-19 pandemic. The main methods used are bibliometric analysis and systematic literature review. Evaluating the number of assets which lead to optimal diversification is not an easy task as it is impacted by a huge number of different factors: the way systematic risk is measured, the investment universe (size, asset classes and features of the asset classes), the investor’s characteristics, the change over time of the asset features, the model adopted to measure diversification (i.e., equally weighted versus optimal allocation), the frequency of the data that is being used, together with the time horizon, conditions in the market that the study refers to, etc. Our paper provides additional support for the fact that (1) a generalized optimal number of stocks that constitute a well-diversified portfolio does not exist for whichever market, period or investor. Recent studies further suggest that (2) the size of a well-diversified portfolio is larger today than in the past, (3) this number is lower in emerging markets compared to developed financial markets, (4) the higher the stock correlations with the market, the lower the number of stocks required for a well-diversified portfolio for individual investors, and (5) machine learning methods could potentially improve the investment decision process. Our results could be helpful to private and institutional investors in constructing and managing their portfolios and provide a framework for future research.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.002 | 0.001 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".