Interrater Agreement between Nurses for the Pediatric Canadian Triage and Acuity Scale in a Tertiary Care Center
Bibliographic record
Abstract
OBJECTIVES: The objective was to measure the interrater agreement between nurses assigning triage levels to children visiting a pediatric emergency departments (EDs) assisted by a computerized version of the Pediatric Canadian Triage and Acuity Scale (PedCTAS). METHODS: This was a prospective cohort study evaluating children triaged from Level 2 (emergent) to Level 5 (nonurgent). A convenience sample of patients triaged during 38 shifts from April to September 2007 in a tertiary care pediatric ED was evaluated. All patients were initially triaged by regular triage nurses using a computerized version of the PedCTAS. Research nurses performed a second evaluation blinded to the first evaluation using the same triage tool. These research nurses were regular ED nurses performing extra hours for research purposes exclusively. The primary outcome measure was the interrater agreement between the two nurses as measured by the linear weighted kappa score. Secondary outcomes included the proportion of patient for which nurses did not apply the triage level suggested by Staturg (override) and agreement for these overrides. RESULTS: A total of 499 patients were recruited. The overall interrater agreement was moderate (linear weighted kappa score of 0.55 [95% confidence interval {CI} = 0.48 to 0.61] and quadratic weighted kappa score of 0.61 [95% CI = 0.42 to 0.80]). There was a discrepancy of more than one level in only 10 patients (2% of the study population). Overrides occurred in 23.2 and 21.8% for regular and research triage nurses, respectively. These overrides were equally distributed between increase and decrease in triage level. CONCLUSIONS: Nurses using Staturg, which is a computerized version of the PedCTAS, demonstrated moderate interrater agreement for assignment of triage level to children presenting to a pediatric ED.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".