Validating the Emergency Department Avoidability Classification (EDAC): A cluster randomized single-blinded agreement study
Bibliographic record
Abstract
INTRODUCTION: The Emergency Department Avoidability Classification (EDAC) retrospectively classifies emergency department (ED) visits that could have been safely managed in subacute primary care settings, but has not been validated against a criterion standard. A validated EDAC could enable accurate and reliable quantification of avoidable ED visits. We compared agreement between the EDAC and ED physician judgements to specify avoidable ED visits. MATERIALS AND METHODS: We conducted a cluster randomized, single-blinded agreement study in an academic hospital in Hamilton, Canada. ED visits between January 1, 2019, and December 31, 2019 were clustered based on EDAC classes and randomly sampled evenly. A total of 160 ED visit charts were randomly assigned to ten participating ED physicians at the academic hospital for evaluation. Physicians judged if the ED visit could have been managed appropriately in subacute primary care (an avoidable visit); each ED visit was evaluated by two physicians independently. We measured interrater agreement between physicians with a Cohen's kappa and 95% confidence intervals (CI). We evaluated the correlation between the EDAC and physician judgements using a Spearman rank correlation and ordinal logistic regression with odds ratios (ORs) and 95% CIs. We examined the EDAC's precision to identify avoidable ED visits using accuracy, sensitivity and specificity. RESULTS: ED physicians agreed on 139 visits (86.9%) with a kappa of 0.69 (95% CI 0.59-0.79), indicating substantial agreement. Physicians judged 96.2% of ED visits classified as avoidable by the EDAC as suitable for management in subacute primary care. We found a high correlation between the EDAC and physician judgements (0.64), as well as a very strong association to classify avoidable ED visits (OR 80.0, 95% CI 17.1-374.9). The EDACs avoidable and potentially avoidable classes demonstrated strong accuracy to identify ED visits suitable for management in subacute care (82.8%, 95% CI 78.2-86.8). DISCUSSION: The EDAC demonstrated strong evidence of criterion validity to classify avoidable ED visits. This classification has important potential for accurately monitoring trends in avoidable ED utilization, measuring proportions of ED volume attributed to avoidable visits and informing interventions intended at reducing ED use by patients who do not require emergency or life-saving healthcare.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.087 | 0.088 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.003 | 0.002 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.002 | 0.004 |
| Scholarly communication | 0.001 | 0.002 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.003 | 0.001 |
| Insufficient payload (model declined to judge) | 0.003 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".