MétaCan
Menu
← Back to cohort
Record W7095031440 · doi:10.5281/zenodo.17446280

Fair AI Assisted Triage in the Emergency Department: A Deployment Framework for a Canadian Teaching Hospital

2025· article· W7095031440 on OpenAlexaboutno aff

Bibliographic record

VenueZenodo (CERN European Organization for Nuclear Research) · 2025
Typearticle
Language
FieldMedicine
TopicArtificial Intelligence in Healthcare and Education
Canadian institutionsnot available
Fundersnot available
KeywordsTriageWorkflowSoftware deploymentEmergency departmentTest (biology)Emergency nursingPatient safety

Abstract

fetched live from OpenAlex

Abstract Emergency departments are crowded, high stakes clinical environments where clinicians need to make fast decisions with incomplete information, which creates real risk for diagnostic error and for delays in care for the sickest patients. [1] Artificial intelligence tools are now able to use triage vital signs, presenting complaint, free text nursing notes and real time operational data to flag people who are at high risk of needing critical care or early intervention, sometimes more accurately than traditional triage scales. [2][3][9] At the same time, there is clear evidence that current triage processes can be inconsistent and inequitable for racialized patients, patients who do not speak English as a primary language, and other structurally marginalized groups. [7] A naive AI system can make that worse if it just learns those same patterns of under triage and delay, and deploys them at scale. [4][7] The goal in this paper is to translate the recent evidence base into something pragmatic: a stepwise framework that a Canadian academic emergency department using the Canadian Triage and Acuity Scale (CTAS) could realistically follow to introduce an AI assisted triage decision support tool in a way that is clinically useful, auditable, and explicitly built to protect equity. [1][2][3][4][5][6][7][8][9] I call this the FAIR ED Triage Framework, and it has eight parts: (1) define the clinical problem and equity gap, (2) lock down data governance and consent, (3) train explainable models on local data, (4) stress test bias and validate prospectively, (5) integrate the tool into nurse and physician workflow instead of replacing human judgment, (6) monitor safety and drift in real time, (7) communicate openly with patients and communities, and (8) align oversight with national triage policy and World Health Organization guidance on responsible AI in health. [1][2][3][4][5][6][7][8][9] Doing these eight steps up front turns AI triage from “cool model with a good AUROC” into a governed clinical safety intervention that an emergency department chief and a hospital ethics board can defend publicly. [2][3][4][5][6][7][8][9] Keywords: emergency department triage; artificial intelligence; equity; CTAS; decision support; diagnostic safety. [1][2][3][4][5][6][7][8][9] 1. Introduction Triage is the front door of the emergency department and it quietly controls who gets seen first, who waits, and in practice who is exposed to the most clinical risk. [2][6][7] In Canada, triage nurses assign a CTAS score from Level 1 (Resuscitation) to Level 5 (Non Urgent) to rank urgency and trigger time to physician targets, and CTAS is a national standard that is updated through formal expert guidelines. [6] In the United States and many other places, a similar role is played by the Emergency Severity Index (ESI), which is partly subjective and tries to capture both acuity and expected resource use. [2][7] A major problem is that nurses are doing this under extreme cognitive load (noise, time pressure, multiple patients at once, constant interruptions) and often before any diagnostics are back, so risk stratification is difficult and highly operator dependent. [1][2][7] That pressure to decide fast with incomplete data is the same pressure that creates diagnostic error downstream in the emergency department. [1] The result in the real world is that triage is not only variable but in some cases systematically unfair. [7] A multicenter study of almost 250,000 adult visits across seven academic and community emergency departments in a US health system found that patients who identified as Black, Hispanic, or Other race and ethnicity were frequently assigned less acute ESI scores than White patients despite having the same chief symptom such as chest pain or abdominal pain, and those same patients then went on to receive more intense physician workups, which strongly suggests they were actually sicker than the score implied. [7] The same study showed that patients whose primary language was not English were also more likely to be assigned less acute triage scores and still ended up needing more workup, which means language itself was operating like a barrier at triage. [7] When you remember that triage score determines how long someone waits in the waiting room, those findings are not just theoretical bias, they are differences in access that can translate directly into harm. [7] At the same time, emergency medicine is now seeing a wave of machine learning tools that claim to improve triage. [2][3][9] One of the best studied examples is an electronic triage system that uses machine learning on triage vitals, chief complaint, and past medical history from the electronic health record to predict immediate clinical risk such as critical care needs, emergency procedures, or hospital admission, instead of just guessing resource use. [2] In a large multi-hospital dataset, those predictions reached areas under the curve between about 0.73 and 0.92 for outcomes like “critical care within hours of arrival,” and the system outperformed the Emergency Severity Index in surfacing high-risk patients hiding in the giant middle acuity bucket. [2] Systematic review work that pooled 60 studies also found that models built with techniques like gradient boosted trees, deep neural networks, and natural language processing of triage nursing notes can outperform traditional judgment or rule-based systems on targets like predicting ICU admission, early mortality, or need for rapid intervention. [3] Canadian emergency departments are already under extreme flow pressure from boarding and crowding, so anything that can identify truly sick patients faster and route lower acuity patients more safely is attractive from both a patient safety and an operations perspective. [1][3][6][8] However, machine learning in triage can recreate and automate the same inequities we already see. [4][7] A 2025 PLOS Digital Health study trained an extreme gradient boosting model (XGBoost) on more than 170,000 emergency department visits, all triaged as ESI level 3 (the “everything is urgent but not dying” group), to predict which patients would face prolonged waits of at least 30 minutes. [4] The model achieved an area under the receiver operating characteristic curve around 0.81 and did a solid job operationally, but when the team evaluated fairness, they found that false negative rates and other error metrics differed across sex, race and ethnicity, and insurance status, which means some demographic groups were less likely to get flagged for long waits even when they actually would wait. [4] The authors explicitly argue that fairness testing has to happen before deployment or else you risk encoding structural access gaps into the digital layer. [4] That is the exact same ethical issue we just saw in human triage. [4][7] Global policy guidance has now caught up to this concern. [5] The World Health Organization released guidance in 2025 on large multimodal models in health that says health AI tools should not be deployed without clear human oversight, documentation of limitations, transparency about training data, active monitoring for harm, and named accountability for decisions. [5] This guidance treats bias, equity, and transparency as clinical safety issues, not as optional ethics talking points. [5] Canadian triage culture already has a national structure in CTAS that is explicitly described as a patient safety and quality improvement tool, with defined escalation pathways for higher acuity levels. [6] This creates a natural place to anchor AI assisted triage and fairness auditing inside the existing Canadian standard instead of inventing a parallel system. [1][5][6] In the rest of this paper, I build a practical framework for safe, fair, AI assisted emergency triage that a Canadian academic emergency department could actually implement while staying aligned with CTAS and the World Health Organization. [1][2][3][4][5][6][7][8][9] 2. Approach and Scope This is a narrative translational review and design proposal, not a single site model training study. [1][2][3][4][5][6][7][8][9] I pulled concepts, quantitative findings, and governance recommendations from recent peer reviewed emergency medicine and informatics work on triage automation, diagnostic safety, bias, and emergency department operations. [1][2][3][4][5][6][7][8][9] These included machine learning based electronic triage systems that outperform ESI in early risk recognition, systematic reviews of triage prediction models, fairness audits of triage and wait time prediction, updated CTAS guidelines, and World Health Organization guidance on AI governance. [1][2][3][4][5][6][7][8][9] I focused on literature from about 2018 to 2025 because that is when real time, EHR embedded triage AI moved from proof of concept to live or near-live clinical pilots, including models that combine structured triage vitals with nurse free text through deep attention or transformer style architectures. [2][3][5][8][9] I then translated those findings into a step by step deployment framework for a Canadian teaching hospital emergency department, which runs CTAS in daily practice and faces the same crowding, throughput pressure, and equity concerns that show up in the international literature. [1][3][4][6][7][8] The main output is the FAIR ED Triage Framework, which is meant to function like a checklist for clinical leaders, informatics teams, and ethics boards. [1][2][3][4][5][6][7][8][9] 3. The FAIR ED Triage Framework The FAIR ED Triage Framework has eight steps. [1][2][3][4][5][6][7][8][9] The goal is not to “replace triage nurses with AI.” [1][5][6] The goal is to standardize early risk recognition, reduce cognitive overload at the front door, and actively narrow inequities in time to care. [1][2][3][4][5][6][7][8][9] 3.1 Step 1. Define the clinical problem and the equity gap Before writing a single line of code, the emergenc

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.007
metaresearch head score (Gemma)0.012
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: none
GenreCandidate signal: Methods · Consensus signal: none
Teacher disagreement score0.324
Threshold uncertainty score0.651

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0070.012
Meta-epidemiology (narrow)0.0010.000
Meta-epidemiology (broad)0.0000.000
Bibliometrics0.0010.001
Science and technology studies0.0040.003
Scholarly communication0.0030.002
Open science0.0030.004
Research integrity0.0020.001
Insufficient payload (model declined to judge)0.0060.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.100
GPT teacher head0.386
Teacher spread0.286 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designNot applicable
Domainnot available
GenreMethods

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2025
Admission routes1
Has abstractyes

Explore more

Same venueZenodo (CERN European Organization for Nuclear Research)→Same topicArtificial Intelligence in Healthcare and Education→French-language works237,207→