Additional file 1 of Characterizing people with frequent emergency department visits and substance use: a retrospective cohort study of linked administrative data in Ontario, Alberta, and B.C., Canada
Bibliographic record
Abstract
Additional file 1: Supplementary Table 1. Summary of diagnostic codes for substance use-related presentations. List of diagnostic codes used to define substance use categories in study cohort definition. a. Supplementary Table 1.1. Summary of ICD-10-CA codes for substance use categories. List of ICD-10-CA diagnostic codes used to define substance use categories. b. Supplementary Table 1.2. Summary of ICD-9 codes for substance use categories. List of ICD-9 diagnostic codes used to define substance use categories Supplementary Table 2. Variable checklist for each database. Summary of databases used to characterize study cohort (NACRS, DAD, HMHDB, MSP, Pharmanet, Vital Events and Statistics), and summary of variables characterized within each database. Supplementary Table 3. Pseudo F values for number of subgroups for Ontario, from 1-10. Summary of subgroup numbers and associated pseudo F values used for cluster analysis in Ontario. Supplementary Table 4. Pseudo F values for number of subgroups for Alberta, from 1-10. Summary of subgroup numbers and associated pseudo F values used for cluster analysis in Alberta. Supplementary Table 5. Pseudo F values for number of subgroups for B.C., from 1-10. Summary of subgroup numbers and associated pseudo F values used for cluster analysis in B.C. Supplementary Figure 1. Visual Representation of Subgroups for Ontario. Graphical representation of clustering variables used to define subgroups in Ontario. Supplementary Figure 2. Visual Representation of Subgroups for Alberta. Graphical representation of clustering variables used to define subgroups in Alberta. Supplementary Figure 3. Visual Representation of Subgroups for B.C. Graphical representation of clustering variables used to define subgroups in B.C. Supplementary Table 6. Demographic and healthcare utilization characteristics of subgroups of people with frequent ED visits (top 10%) and substance use in Ontario, Alberta, and B.C. (April 1st, 2014 to March 31st, 2015). Detailed summary of subgroups’ demographic and healthcare utilization characteristics to complement main study Tables.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.017 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.003 | 0.009 |
| Science and technology studies | 0.002 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.002 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.418 | 0.025 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".