Sustainable vs. Non-Sustainable Assets: A Deep Learning-Based Dynamic Portfolio Allocation Strategy
Bibliographic record
Abstract
This article aims to investigate the impact of sustainable assets on dynamic portfolio optimization under varying levels of investor risk aversion, particularly during turbulent market conditions. The analysis compares the performance of two portfolio types: (i) portfolios composed of non-sustainable assets such as fossil energy commodities and conventional equity indices, and (ii) mixed portfolios that combine non-sustainable and sustainable assets, including renewable energy, green bonds, and precious metals using advanced Deep Reinforcement Learning models (including TD3 and DDPG) based on risk and transaction cost- sensitive in portfolio optimization against the traditional Mean-Variance model. Results show that incorporating clean and sustainable assets significantly enhances portfolio returns and reduces volatility across all risk aversion profiles. Moreover, the Deep Reinforcing Learning optimization models outperform classical MV optimization, and the RTC-LSTM-TD3 optimization strategy outperforms all others. The RTC-LSTM-TD3 optimization achieves an annual return of 24.18% and a Sharpe ratio of 2.91 in mixed portfolios (sustainable and non-sustainable assets) under low risk aversion (λ = 0.005), compared to a return of only 8.73% and a Sharpe ratio of 0.67 in portfolios excluding sustainable assets. To the best of the authors’ knowledge, this is the first study that employs the DRL framework integrating risk sensitivity and transaction costs to evaluate the diversification benefits of sustainable assets. Findings offer important implications for portfolio managers to leverage the benefits of sustainable diversification, and for policymakers to encourage the integration of sustainable assets, while addressing fiduciary responsibilities.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".