Predictive Models on the Therapeutic Effects of Diverse Music Repertoires on Mental Health
Bibliographic record
Abstract
Music therapy is a recognized field of therapeutic intervention used in a diverse variety of contexts, including hospital settings, retirement homes, palliative care, and community programs. It leverages intrinsic qualities of music to address mental health issues and enhance overall mental health. A challenge for music therapists working in hospice is lack of suitable musical repertoires, called “music unpreparedness”. This study aims to use a machine learning model to evaluate repertoires and predict their effects on an individual’s mental health given their age and mental health condition, if any. Streaming services such as Spotify can also implement this model to improve the overall effects of services on listeners’ mental health–for instance, promoting playlists that will likely improve, rather than worsen, the user’s mental health on their home page. To train the model, an open-access dataset was used containing features such as favorite genre, age, and mental health condition. Categorical variables including ‘Fav genre’ and ‘Frequency [genre]’ were prepared using one-hot encoding and the dataset was split into training and testing sets, using a stratified approach to maintain the distribution of mental health conditions and music preferences. A decision tree model was chosen due to its interpretability and ability to handle categorical data. The model was trained with the entropy criterion and a maximum depth of 6, achieving a training accuracy of 82% and a validation accuracy of 75%. Key features influencing the model included listening hours/frequency, favorite genres, and mental health condition levels. The model provides insights into how these factors contribute to the perceived effect of music on mental health, offering valuable predictions that can enhance therapeutic interventions and user experiences on streaming platforms. Future improvements include incorporating more diverse datasets to improve the model's generalizability. Additionally, other machine learning algorithms can be explored to enhance prediction accuracy.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.003 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".