A Systematic Review of Instruments to Identify Mental Health and Substance Use Problems Among Children in the Emergency Department
Bibliographic record
Abstract
OBJECTIVE: Specialized instruments to screen and diagnose mental health problems in children and adolescents are not yet standard components of clinical assessments in emergency departments (EDs). We conducted a systematic review to investigate the psychometric properties, accuracy, and performance metrics of instruments used in the ED to identify pediatric mental health and substance use problems. METHODS: We searched seven electronic databases and the gray literature for psychometric validation studies, diagnostic studies, and cohort studies that assessed any instrument to screen for or diagnose mental illness, emotional or behavioral problems, or substance use disorders. Studies had to include children and adolescents with mental health presentations or positive screens for substance use. Two reviewers independently screened studies for relevance and quality. Diagnostic study quality was assessed with the four QUADAS-2 domains. Psychometric study quality was assessed with published criteria for instrument reliability, validity, and usability. We present a descriptive analysis of the reported psychometric properties and diagnostic performance of instruments for each study. RESULTS: Of the 4,832 references screened, 14 met inclusion criteria. Included studies evaluate 18 instruments for identifying suicide risk (six studies), alcohol use disorders (six studies), mood disorders (one study), and ED decision making (need for assessment, admission; one study). Nine studies include a psychometric focus but quality varies, with no studies fully meeting criteria for reliability, validity, and usability. Seven studies examine diagnostic performance of an instrument, but no study has a low risk of bias for all QUADAS-2 domains. The HEADS-ED instrument has good inter-rater reliability (r = 0.785) for identifying general mental health problems and modest evidence for ruling in patients requiring hospital admission (positive likelihood ratio [LR+] = 6.30). Internal consistency (reliability) varies for instruments to screen for suicide risk (α = 0.46-0.97), and no instruments have both high sensitivity and high specificity. The Ask Suicide-Screening Questions (ASQ) is highly sensitive (98%) and has strong evidence for ruling out risk (negative likelihood ratio [LR-] = 0.04). Among screening instruments for alcohol use disorders, internal consistency is high for the consumption subscale of the Alcohol Use Disorders Identification Test (α = 0.83-0.88) and the Adolescent Drinking Index (α = 0.92). Both instruments also had sound internal validity. Diagnostically, a two-item instrument based on DSM-IV criteria is the most accurate in identifying patients with a disorder (area under the curve = 0.89) and has modest evidence for ruling in and out risk (LR+ = 8.80, LR- = 0.13). CONCLUSIONS: From available evidence, we recommend that ED clinicians use 1) the HEADS-ED to rule in ED admission among pediatric patients with visits for mental health care, 2) the ASQ to rule out suicide risk among pediatric patients with any visit type, and 3) the DSM-IV two-item instrument to rule in/rule out alcohol use disorders among pediatric patients currently using alcohol. These instruments require minimal to no training or time commitment. We also recommend that clinicians become familiar with each instrument's psychometric properties to understand the quality of the evidence base. In this review, however, we identify methodologic limitations in the evidence base. To develop a robust evidence base, additional research is necessary.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.001 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.004 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".