Commentary on Afzali <i>et al</i>. (2019): Two data sets are better than one
Notice bibliographique
Résumé
The use of large data sets in addiction research is welcome, because statistical power is increased. When applied to large data sets, machine learning can help with interpreting variable importance and with quantifying reproducibility. However, application of machine learning in the real world requires consideration of several factors, such as economic cost. Afzali et al. employed two methods to identify the most important variables associated with adolescent alcohol use. First, the best-performing machine was one that employed regularized regression using the Elastic Net 2, a method that automatically selects out the most predictive variables in a data set, performing a similar role to stepwise regression but with distinct advantages. The Elastic Net attenuates overfitting when selecting variables, whereas stepwise regression is especially prone to this 3, and Elastic Net regularization accepts or rejects groups of correlated variables: this is important in addiction research, where measurement variables are typically correlated. Secondly, Afzali et al. also systematically included or excluded variables associated with various domains. Notably, for both data sets, inclusion of all domains produced the most accurate predictions. These results provide guidance for variables that we should assay, given time and budgetary constraints, if our goal is to predict alcohol-related behaviour with high accuracy: measuring psychopathology and personality should be a priority, but it is also worth obtaining some data from a wide variety of other domains. Using two large independent data sets, Afzali et al. were able to use external cross-validation to evaluate the performance of their model. External cross-validation directly tests a model's ability to generalize to previously unseen data because the model is trained on one data set and then tested on a separate data set. External cross-validation therefore speaks directly to replication issues in science 4, 5. The use of external cross-validation in Afzali et al.’s work highlights an interesting aspect of this validation method when applied to addiction data. Unlike other data-driven fields (e.g. internet search), addiction researchers cannot easily add more data for validation purposes. Data-driven addiction research is therefore likely to be advanced by interactions among scientists practising ‘team science’, with research distributed throughout sites to increase statistical power and for testing generalizability 6. Variables do not have to be identical throughout sites, as was the case in Afzali et al., who used different questions to measure alcohol use in the Australian and Canadian samples. Indeed, the external validation of models despite the use of slightly different variables is a strength—scientific findings should be robust to reasonable deviations in methodology between sites. An avenue for future research should be to validate Afzali et al.’s findings in additional samples, particularly in a wider range of cultures. Afzali et al.’s results tell us what the most important predictors of adolescent alcohol use are, but there is a sizable gap between research-focused models and their application on a population level 7. For example, economic analyses are needed to quantify the relative efficacy of machine learning versus traditional methods to identify adolescents at high risk of alcohol initiation. Machine learning in the wild must accommodate a host of other factors not typically examined in a research setting. For example, the cost of misclassification depends on the nature of any subsequent intervention. If false positives (incorrectly classifying as high risk) result in allocation to a resource-intensive intervention programme, then the machine should be trained to avoid false positives. Alternatively, if the intervention is low-cost and benign (e.g. delivering information online), then the machine should avoid false negatives: it is better to intervene for someone at low risk rather than miss someone at high risk. Furthermore, not all variables cost the same to obtain. It is plausible that a weaker, but cheaper, predictor could be more cost-effective than a stronger, more expensive, one at the population level. Afzali et al. have made a valuable contribution to the addiction literature—a reproducible set of findings produced by the combination of sophisticated methods and cross-country collaboration. We are still some distance away from the era of personalized interventions at the earliest stage of substance use disorder, but studies such as those by Afzali et al. certainly represent encouraging initial steps. None.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction distillée sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.
Scores Codex et Gemma par catégorie
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,001 | 0,000 |
| Méta-épidémiologie (sens strict) | 0,000 | 0,000 |
| Méta-épidémiologie (sens large) | 0,000 | 0,000 |
| Bibliométrie | 0,000 | 0,000 |
| Études des sciences et des technologies | 0,000 | 0,000 |
| Communication savante | 0,000 | 0,000 |
| Science ouverte | 0,001 | 0,001 |
| Intégrité de la recherche | 0,001 | 0,004 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,005 | 0,010 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; les deux têtes enseignantes s’accordent sur ce qui est montré ici.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».