Commentary on Afzali <i>et al</i>. (2019): Two data sets are better than one
Bibliographic record
Abstract
The use of large data sets in addiction research is welcome, because statistical power is increased. When applied to large data sets, machine learning can help with interpreting variable importance and with quantifying reproducibility. However, application of machine learning in the real world requires consideration of several factors, such as economic cost. Afzali et al. employed two methods to identify the most important variables associated with adolescent alcohol use. First, the best-performing machine was one that employed regularized regression using the Elastic Net 2, a method that automatically selects out the most predictive variables in a data set, performing a similar role to stepwise regression but with distinct advantages. The Elastic Net attenuates overfitting when selecting variables, whereas stepwise regression is especially prone to this 3, and Elastic Net regularization accepts or rejects groups of correlated variables: this is important in addiction research, where measurement variables are typically correlated. Secondly, Afzali et al. also systematically included or excluded variables associated with various domains. Notably, for both data sets, inclusion of all domains produced the most accurate predictions. These results provide guidance for variables that we should assay, given time and budgetary constraints, if our goal is to predict alcohol-related behaviour with high accuracy: measuring psychopathology and personality should be a priority, but it is also worth obtaining some data from a wide variety of other domains. Using two large independent data sets, Afzali et al. were able to use external cross-validation to evaluate the performance of their model. External cross-validation directly tests a model's ability to generalize to previously unseen data because the model is trained on one data set and then tested on a separate data set. External cross-validation therefore speaks directly to replication issues in science 4, 5. The use of external cross-validation in Afzali et al.’s work highlights an interesting aspect of this validation method when applied to addiction data. Unlike other data-driven fields (e.g. internet search), addiction researchers cannot easily add more data for validation purposes. Data-driven addiction research is therefore likely to be advanced by interactions among scientists practising ‘team science’, with research distributed throughout sites to increase statistical power and for testing generalizability 6. Variables do not have to be identical throughout sites, as was the case in Afzali et al., who used different questions to measure alcohol use in the Australian and Canadian samples. Indeed, the external validation of models despite the use of slightly different variables is a strength—scientific findings should be robust to reasonable deviations in methodology between sites. An avenue for future research should be to validate Afzali et al.’s findings in additional samples, particularly in a wider range of cultures. Afzali et al.’s results tell us what the most important predictors of adolescent alcohol use are, but there is a sizable gap between research-focused models and their application on a population level 7. For example, economic analyses are needed to quantify the relative efficacy of machine learning versus traditional methods to identify adolescents at high risk of alcohol initiation. Machine learning in the wild must accommodate a host of other factors not typically examined in a research setting. For example, the cost of misclassification depends on the nature of any subsequent intervention. If false positives (incorrectly classifying as high risk) result in allocation to a resource-intensive intervention programme, then the machine should be trained to avoid false positives. Alternatively, if the intervention is low-cost and benign (e.g. delivering information online), then the machine should avoid false negatives: it is better to intervene for someone at low risk rather than miss someone at high risk. Furthermore, not all variables cost the same to obtain. It is plausible that a weaker, but cheaper, predictor could be more cost-effective than a stronger, more expensive, one at the population level. Afzali et al. have made a valuable contribution to the addiction literature—a reproducible set of findings produced by the combination of sophisticated methods and cross-country collaboration. We are still some distance away from the era of personalized interventions at the earliest stage of substance use disorder, but studies such as those by Afzali et al. certainly represent encouraging initial steps. None.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.004 |
| Insufficient payload (model declined to judge) | 0.005 | 0.010 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; both teacher heads agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".