A Three-Tier Artificial Intelligence Model for COVID-19 Triage Using Pharyngeal Images (Preprint)
Notice bibliographique
Résumé
Background: SARS-CoV-2 remains a common cause of acute respiratory illness; however, symptom-based triage poorly discriminates it from other febrile conditions. A recently developed artificial intelligence (AI)-powered pharyngeal camera acquires pharyngeal images and clinical data to assist in influenza diagnosis; leveraging this workflow, we evaluated an adjunct AI algorithm (COVID-19-AI) that reports high, medium, or low suspicion to guide whether SARS-CoV-2 testing should subsequently be performed. Objective: This study aimed to report diagnostic accuracy outcomes and clinical utility of the COVID-19-AI as a triage support tool. Methods: We conducted a performance evaluation using a prospectively collected multicenter dataset from 26 Japanese institutions between December 2023 and March 2024. Patients with suspected influenza or COVID-19 were eligible. The COVID-19-AI algorithm, a stacked ensemble of a Swin Transformer and boosting models, was developed using pharyngeal images combined with routine clinical variables from 2133 patients, and it produced a 3-tier output. Classification thresholds were predefined to optimize clinical rule-out and rule-in utilities. Diagnostic performance was assessed in 696 independent patients against centralized reverse transcription polymerase chain reaction-confirmed SARS-CoV-2 infection under 2 prespecified operating criteria: inclusive (high or medium=positive and low=negative) and strict (high=positive and medium or low=negative). A subanalysis stratified accuracy by time from symptom onset (12-hour bins to 72 hours). Results: Among 696 analyzed participants (all Asian), 247 (35.5%) had reverse transcription polymerase chain reaction-confirmed SARS-CoV-2 infection. The COVID-19-AI categorized 12.4% (n=86), 72.8% (n=507), and 14.8% (n=103) patients as high, medium, and low suspicion, respectively. Under the inclusive criteria, sensitivity of COVID-19-AI was 93.9% (95% CI 90.4%-96.4%), specificity was 19.6% (95% CI 16.1%-23.5%), and negative predictive value was 85.4% (95% CI 77.6%-91.3%). Under the strict criteria, sensitivity was 24.7% (95% CI 19.6%-30.4%), specificity was 94.4% (95% CI 92.0%-96.3%), and positive predictive value was 70.9% (95% CI 60.7%-79.8%). Across 12-hour onset strata, sensitivity under the inclusive criteria remained ≥92.0% and specificity under the strict criteria remained ≥83.3%; no pronounced temporal trend was observed. Additionally, an integrated model using both pharyngeal images and clinical variables (area under the receiver operating characteristic curve [AUROC] 0.78) outperformed models using only clinical variables (AUROC 0.75) or images alone (AUROC 0.71); feature importance analysis further confirmed that pharyngeal image information was the most influential individual predictor, providing greater predictive value than any single clinical variable. Conclusions: Embedded within AI-powered pharyngeal camera workflows, a 3-tier AI suspicion output enables complementary operating behaviors-high sensitivity to rule out COVID-19 (inclusive criteria) and high specificity to support immediate infection control measures (strict criteria), although these criteria involve inherent trade-offs with low specificity and low sensitivity, respectively. Performance stability across onset times suggests robustness to symptom chronology, offering a standardized tool for clinical triage.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,002 | 0,005 |
| Méta-épidémiologie (sens strict) | 0,001 | 0,001 |
| Méta-épidémiologie (sens large) | 0,001 | 0,001 |
| Bibliométrie | 0,001 | 0,001 |
| Études des sciences et des technologies | 0,001 | 0,001 |
| Communication savante | 0,003 | 0,002 |
| Science ouverte | 0,002 | 0,001 |
| Intégrité de la recherche | 0,002 | 0,001 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,008 | 0,002 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».