MétaCan
Menu
← Retour à la cohorte
Enregistrement W7095031440 · doi:10.5281/zenodo.17446280

Fair AI Assisted Triage in the Emergency Department: A Deployment Framework for a Canadian Teaching Hospital

2025· article· W7095031440 sur OpenAlexaboutno aff

Notice bibliographique

RevueZenodo (CERN European Organization for Nuclear Research) · 2025
Typearticle
Langue
DomaineMedicine
ThématiqueArtificial Intelligence in Healthcare and Education
Établissements canadiensnon disponible
Organismes subventionnairesnon disponible
Mots-clésTriageWorkflowSoftware deploymentEmergency departmentTest (biology)Emergency nursingPatient safety

Résumé

récupéré en direct d'OpenAlex

Abstract Emergency departments are crowded, high stakes clinical environments where clinicians need to make fast decisions with incomplete information, which creates real risk for diagnostic error and for delays in care for the sickest patients. [1] Artificial intelligence tools are now able to use triage vital signs, presenting complaint, free text nursing notes and real time operational data to flag people who are at high risk of needing critical care or early intervention, sometimes more accurately than traditional triage scales. [2][3][9] At the same time, there is clear evidence that current triage processes can be inconsistent and inequitable for racialized patients, patients who do not speak English as a primary language, and other structurally marginalized groups. [7] A naive AI system can make that worse if it just learns those same patterns of under triage and delay, and deploys them at scale. [4][7] The goal in this paper is to translate the recent evidence base into something pragmatic: a stepwise framework that a Canadian academic emergency department using the Canadian Triage and Acuity Scale (CTAS) could realistically follow to introduce an AI assisted triage decision support tool in a way that is clinically useful, auditable, and explicitly built to protect equity. [1][2][3][4][5][6][7][8][9] I call this the FAIR ED Triage Framework, and it has eight parts: (1) define the clinical problem and equity gap, (2) lock down data governance and consent, (3) train explainable models on local data, (4) stress test bias and validate prospectively, (5) integrate the tool into nurse and physician workflow instead of replacing human judgment, (6) monitor safety and drift in real time, (7) communicate openly with patients and communities, and (8) align oversight with national triage policy and World Health Organization guidance on responsible AI in health. [1][2][3][4][5][6][7][8][9] Doing these eight steps up front turns AI triage from “cool model with a good AUROC” into a governed clinical safety intervention that an emergency department chief and a hospital ethics board can defend publicly. [2][3][4][5][6][7][8][9] Keywords: emergency department triage; artificial intelligence; equity; CTAS; decision support; diagnostic safety. [1][2][3][4][5][6][7][8][9] 1. Introduction Triage is the front door of the emergency department and it quietly controls who gets seen first, who waits, and in practice who is exposed to the most clinical risk. [2][6][7] In Canada, triage nurses assign a CTAS score from Level 1 (Resuscitation) to Level 5 (Non Urgent) to rank urgency and trigger time to physician targets, and CTAS is a national standard that is updated through formal expert guidelines. [6] In the United States and many other places, a similar role is played by the Emergency Severity Index (ESI), which is partly subjective and tries to capture both acuity and expected resource use. [2][7] A major problem is that nurses are doing this under extreme cognitive load (noise, time pressure, multiple patients at once, constant interruptions) and often before any diagnostics are back, so risk stratification is difficult and highly operator dependent. [1][2][7] That pressure to decide fast with incomplete data is the same pressure that creates diagnostic error downstream in the emergency department. [1] The result in the real world is that triage is not only variable but in some cases systematically unfair. [7] A multicenter study of almost 250,000 adult visits across seven academic and community emergency departments in a US health system found that patients who identified as Black, Hispanic, or Other race and ethnicity were frequently assigned less acute ESI scores than White patients despite having the same chief symptom such as chest pain or abdominal pain, and those same patients then went on to receive more intense physician workups, which strongly suggests they were actually sicker than the score implied. [7] The same study showed that patients whose primary language was not English were also more likely to be assigned less acute triage scores and still ended up needing more workup, which means language itself was operating like a barrier at triage. [7] When you remember that triage score determines how long someone waits in the waiting room, those findings are not just theoretical bias, they are differences in access that can translate directly into harm. [7] At the same time, emergency medicine is now seeing a wave of machine learning tools that claim to improve triage. [2][3][9] One of the best studied examples is an electronic triage system that uses machine learning on triage vitals, chief complaint, and past medical history from the electronic health record to predict immediate clinical risk such as critical care needs, emergency procedures, or hospital admission, instead of just guessing resource use. [2] In a large multi-hospital dataset, those predictions reached areas under the curve between about 0.73 and 0.92 for outcomes like “critical care within hours of arrival,” and the system outperformed the Emergency Severity Index in surfacing high-risk patients hiding in the giant middle acuity bucket. [2] Systematic review work that pooled 60 studies also found that models built with techniques like gradient boosted trees, deep neural networks, and natural language processing of triage nursing notes can outperform traditional judgment or rule-based systems on targets like predicting ICU admission, early mortality, or need for rapid intervention. [3] Canadian emergency departments are already under extreme flow pressure from boarding and crowding, so anything that can identify truly sick patients faster and route lower acuity patients more safely is attractive from both a patient safety and an operations perspective. [1][3][6][8] However, machine learning in triage can recreate and automate the same inequities we already see. [4][7] A 2025 PLOS Digital Health study trained an extreme gradient boosting model (XGBoost) on more than 170,000 emergency department visits, all triaged as ESI level 3 (the “everything is urgent but not dying” group), to predict which patients would face prolonged waits of at least 30 minutes. [4] The model achieved an area under the receiver operating characteristic curve around 0.81 and did a solid job operationally, but when the team evaluated fairness, they found that false negative rates and other error metrics differed across sex, race and ethnicity, and insurance status, which means some demographic groups were less likely to get flagged for long waits even when they actually would wait. [4] The authors explicitly argue that fairness testing has to happen before deployment or else you risk encoding structural access gaps into the digital layer. [4] That is the exact same ethical issue we just saw in human triage. [4][7] Global policy guidance has now caught up to this concern. [5] The World Health Organization released guidance in 2025 on large multimodal models in health that says health AI tools should not be deployed without clear human oversight, documentation of limitations, transparency about training data, active monitoring for harm, and named accountability for decisions. [5] This guidance treats bias, equity, and transparency as clinical safety issues, not as optional ethics talking points. [5] Canadian triage culture already has a national structure in CTAS that is explicitly described as a patient safety and quality improvement tool, with defined escalation pathways for higher acuity levels. [6] This creates a natural place to anchor AI assisted triage and fairness auditing inside the existing Canadian standard instead of inventing a parallel system. [1][5][6] In the rest of this paper, I build a practical framework for safe, fair, AI assisted emergency triage that a Canadian academic emergency department could actually implement while staying aligned with CTAS and the World Health Organization. [1][2][3][4][5][6][7][8][9] 2. Approach and Scope This is a narrative translational review and design proposal, not a single site model training study. [1][2][3][4][5][6][7][8][9] I pulled concepts, quantitative findings, and governance recommendations from recent peer reviewed emergency medicine and informatics work on triage automation, diagnostic safety, bias, and emergency department operations. [1][2][3][4][5][6][7][8][9] These included machine learning based electronic triage systems that outperform ESI in early risk recognition, systematic reviews of triage prediction models, fairness audits of triage and wait time prediction, updated CTAS guidelines, and World Health Organization guidance on AI governance. [1][2][3][4][5][6][7][8][9] I focused on literature from about 2018 to 2025 because that is when real time, EHR embedded triage AI moved from proof of concept to live or near-live clinical pilots, including models that combine structured triage vitals with nurse free text through deep attention or transformer style architectures. [2][3][5][8][9] I then translated those findings into a step by step deployment framework for a Canadian teaching hospital emergency department, which runs CTAS in daily practice and faces the same crowding, throughput pressure, and equity concerns that show up in the international literature. [1][3][4][6][7][8] The main output is the FAIR ED Triage Framework, which is meant to function like a checklist for clinical leaders, informatics teams, and ethics boards. [1][2][3][4][5][6][7][8][9] 3. The FAIR ED Triage Framework The FAIR ED Triage Framework has eight steps. [1][2][3][4][5][6][7][8][9] The goal is not to “replace triage nurses with AI.” [1][5][6] The goal is to standardize early risk recognition, reduce cognitive overload at the front door, and actively narrow inequities in time to care. [1][2][3][4][5][6][7][8][9] 3.1 Step 1. Define the clinical problem and the equity gap Before writing a single line of code, the emergenc

Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.

Comment cette classification a été obtenuedéplier

Prédiction machine sur la base complète

Imitation des enseignants

Ni prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.

score de la tête « metaresearch » (Codex)0,007
score de la tête « metaresearch » (Gemma)0,012
Version: metacan-v3-hybrid-931329e0061cStatut de validation: machine_predicted_unvalidated
Catégories candidatesaucune
Catégories consensuellesaucune
DomaineSignal candidat: aucune · Signal consensuel: aucune
Devis d'étudeSignal candidat: Sans objet · Signal consensuel: aucune
GenreSignal candidat: Méthodes · Signal consensuel: aucune
Score de désaccord entre enseignants0,324
Score d'incertitude au seuil0,651

Scores du classifieur distillé par catégorie (deux têtes)

CatégorieCodexGemma
Métarecherche0,0070,012
Méta-épidémiologie (sens strict)0,0010,000
Méta-épidémiologie (sens large)0,0000,000
Bibliométrie0,0010,001
Études des sciences et des technologies0,0040,003
Communication savante0,0030,002
Science ouverte0,0030,004
Intégrité de la recherche0,0020,001
Charge utile insuffisante (le modèle a refusé de juger)0,0060,001

Scores machine (provisoires)

Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.

Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.

Tête enseignante Opus0,100
Tête enseignante GPT0,386
Écart entre enseignants0,286 · la distance entre les deux têtes enseignantes sur ce seul travail
Statut de validationscore_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découle

Classification

machine, non validée

Prédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.

Les modèles n’ont appliqué aucune catégorie : rien dans la taxonomie ne correspondait à ce travail.
Devis d'étudeSans objet
Domainenon disponible
GenreMéthodes

Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».

En bref

Citations0
Publié2025
Routes d'admission1
Résumé présentoui

Explorer davantage

Même revueZenodo (CERN European Organization for Nuclear Research)→Même sujetArtificial Intelligence in Healthcare and Education→Travaux en français237 207→