Non-Basketball Features in the NBA: Machine Learning and SHAP Analysis of Player Contracts and Team Priorities
Bibliographic record
Abstract
The National Basketball Association (NBA) is the world’s biggest basketball league, with 30 teams across the United States and Canada, which generates billions in annual revenue. An NBA players impact is usually determined based on performance on the court. However, off-court factors influence aspects of the NBA, such as player contracts. This study investigates how non-basketball features, such as off-court factors like Instagram followers and player loyalty, affect player contracts. By using various Machine Learning models, this study analyzed the model’s predictive results trained on a dataset containing only basketball features and compared the results on a dataset containing both basketball and non-basketball features. Furthermore, this study was assisted by the Explainable AI model SHAP to examine how the most valuable teams versus the least valuable teams in the NBA prioritized these non-basketball features. SHAP’s reliability was also assessed for this specific problem. The results showed that incorporating non-basketball features significantly improves the predictive performance of many Machine Learning models, but not Deep Learning models performance in this study. The SHAP analysis revealed that there are differences between highly valuable teams and low-value teams. Highly valuable teams pay for every feature on average more than low-value teams, and if a player were an All Star, it is more likely that this player will be paid more on a highly valuable team. The SHAP assessment test demonstrated its functionality in this case. However, in a general context, SHAP reliability cannot be proved in this study. These results highlight the role of non- basketball features in NBA salaries and offer insights into the application of explainable AI in salary prediction.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".