A comparison of a formal triage scoring system and a quick-look triage approach
Bibliographic record
Abstract
BACKGROUND: Emergency Department (ED) triage systems have become increasingly comprehensive over time, requiring ever more resources such as nursing time and computer support. There are very few studies that have looked at whether this increased complexity results in improved performance. OBJECTIVES: This study looked at one aspect of performance, comparing reliability of triage nurses' (TNs) triage scores utilizing a simple quick-look method with a commonly used, resource-intense, five-level triage system. METHODS: This observational study of TNs was carried out in two urban tertiary-care hospital EDs, in real time, assessing patients arriving consecutively. Immediately upon patients' arrival, TNs were asked to assign triage scores based simply on their observation of the patient and the chief complaint. The patient was then triaged in the department's usual way, utilizing a computer-assisted five-level triage system [Canadian Triage and Acuity Scale (CTAS)]. Agreement between scores was quantified. κ scores were calculated, and weighted by the CTAS score. RESULTS: A total of 496 triage assessments were included. Percent agreement between the quick-look method and the standard CTAS method was 84.5%. κ scores were moderately high. Fourteen patients (2.6%), ultimately classified as CTAS 1 or 2, initially received lower scores from TNs using the quick-look method. No comparison of validity was assessed. CONCLUSION: TNs assigning triage scores to ED patients on arrival, using only chief complaint and observation, were statistically comparable to scores assigned utilizing a resource-intense, comprehensive triage system, but clinically significant discrepancies were identified.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".