Assessing the reliability of pediatric emergency medicine billing code assignment for future consideration as a proxy workload measure
Bibliographic record
Abstract
OBJECTIVES: Prediction of pediatric emergency department (PED) workload can allow for optimized allocation of resources to improve patient care and reduce physician burnout. A measure of PED workload is thus required, but to date no variable has been consistently used or could be validated against for this purpose. Billing codes, a variable assigned by physicians to reflect the complexity of medical decision making, have the potential to be a proxy measure of PED workload but must be assessed for reliability. In this study, we investigated how reliably billing codes are assigned by PED physicians, and factors that affect the inter-rater reliability of billing code assignment. METHODS: A retrospective cross-sectional study was completed to determine the reliability of billing code assigned by physicians (n = 150) at a quaternary-level PED between January 2018 and December 2018. Clinical visit information was extracted from health records and presented to a billing auditor, who independently assigned a billing code-considered as the criterion standard. Inter-rater reliability was calculated to assess agreement between the physician-assigned versus billing auditor-assigned billing codes. Unadjusted and adjusted logistic regression models were used to assess the association between covariables of interest and inter-rater reliability. RESULTS: Overall, we found substantial inter-rater reliability (AC2 0.72 [95% CI 0.64-0.8]) between the billing codes assigned by physicians compared to those assigned by the billing auditor. Adjusted logistic regression models controlling for Pediatric Canadian Triage and Acuity scores, disposition, and time of day suggest that clinical trainee involvement is significantly associated with increased inter-rater reliability. CONCLUSIONS: Our work identified that there is substantial agreement between PED physician and a billing auditor assigned billing codes, and thus are reliably assigned by PED physicians. This is a crucial step in validating billing codes as a potential proxy measure of pediatric emergency physician workload.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.035 | 0.117 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".