Bibliographic record
Abstract
This study delves into predicting song popularity on Spotify by analyzing a dataset of song features from 1986 to 2022. Using linear regression, this paper examines the influence of audio characteristics such as energy, danceability, speechiness, duration, and mode, alongside the year of release. The findings indicate that danceability, more recent release years, and longer track duration are positively associated with higher popularity levels. Conversely, songs in minor keys are more favored than those in major keys. These results highlight the significance of both intrinsic musical qualities and evolving listener preferences over time. The model's robustness is ensured through comprehensive diagnostic tests that validate the assumptions of linearity, normality, and homoscedasticity, confirming the predictive reliability of the identified factors. This research not only enhances the understanding of the dynamics driving music popularity but also provides valuable insights for artists and producers aiming to optimize their music for digital platforms. By focusing on the critical elements that resonate with contemporary audiences, stakeholders can better strategize their music releases to maximize listener engagement and success on streaming platforms.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.002 | 0.002 |
| Open science | 0.002 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".