How well do DSM-5 criteria measure alcohol use disorder in the general population of older Swedish adolescents? An item response theory analysis
Bibliographic record
Abstract
BACKGROUND: This study assesses the psychometric properties of DSM-5 criteria of AUD in older Swedish adolescents using item response theory models, focusing specifically on the precision of the scale at the cut-offs for mild, moderate, and severe AUD. METHODS: Data from the second wave of Futura01 was used. Futura01 is a nationally representative cohort study of Swedish people born 2001 and data for the second wave was collected when participants were 17/18 years old. This study included only participants who had consumed alcohol during the past 12 months (n = 2648). AUD was measured with 11 binary items. A 2-parameter logistic item response theory model (2PL) estimated the items' difficulty and discrimination parameters. RESULTS: 31.8% of the participants met criteria for AUD. Among these, 75.6% had mild AUD, 18.3% had moderate, and 6.1% had severe AUD. A unidimensional AUD model had a good fit and 2PL models showed that the scale measured AUD over all three cut-offs for AUD severity. Although discrimination parameters ranged from moderate (1.24) to very high (2.38), the more commonly endorsed items discriminated less well than the more difficult items, as also reflected in less precision of the estimates at lower levels of AUD severity. The diagnostic uncertainty was pronounced at the cut-off for mild AUD. CONCLUSION: DSM-5 criteria measure AUD with better precision at higher levels of AUD severity than at lower levels. As most older adolescents who fulfil an AUD diagnosis are in the mild category, notable uncertainties are involved when an AUD diagnosis is set in this group.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".