Low-Cost Automated Microscopy and Morphology-Based Machine Learning Classification of Sickle Cell Disease and Beta-Thalassemia in Nepal and Canada
Notice bibliographique
Résumé
INTRODUCTION Sickle cell disease (SCD) is the most common inherited blood disorder, and is associated with a high mortality rate for children in low-resource settings, where the disease burden is high. Early screening and treatment options reduce morbidity and mortality. In addition to screening for SCD, screening for asymptomatic, heterozygous conditions, such as sickle cell trait (SCT) and β-thalassemia trait, is required for effective SCD and thalassemia screening programs. However, rural and remote communities lack accurate low-cost screening techniques. The goal of this study is to augment a common low-cost screening test, the sickling test, using automated microscopy and machine learning to classify individuals with SCD (HbS/β-thalassemia and HbSS), trait conditions (HbAS and HbA/β-thalassemia), and those without common beta-globin disorders (HbAA) for implementation in Nepal. As HbC mutation is not encountered in Nepal, screening for HbSC or HbAC was not included. Additionally, we aim to evaluate the feasibility of 5 low-cost techniques to supplement/replace the local gold standard test (Hb HPLC), which is inaccessible in rural/remote settings. METHODS In this study (ClinicalTrials.gov Identifier: NCT05506358), blood samples were collected and tested at respective clinical sites in Nepal (Mount Sagarmatha Polyclinic and Diagnostic Center, Nepalgunj) and Canada (St. Paul's Hospital and BC Children's Hospital, Vancouver). Participants (138 total: 111 Nepal; 27 Canada) between ages 2 to 74 whose diagnoses were previously established by Hb HPLC were recruited. Five low-cost tests (our augmented sickling test, HbS solubility test, HemoTypeSC, Sickle SCAN, Gazelle Hb variant test) and Hb HPLC were performed on all participants (30 HbAA; 23 HbA/β-thalassemia; 45 HbAS; 11 HbS/β-thalassemia; 29 HbSS). Informed consent was provided by the participants or parents, according to protocols approved by institutional/national research ethics boards. For the sickling test, a sealed wet preparation of blood mixed with 2% sodium metabisulphite was imaged with a robotic microscope (Octopi) with automated scanning/focusing. The images of blood cells were segmented using Cellpose 2.0, allowing for morphological characterization of cells and morphology based classification, with an 80:20 participant-wise split of training and testing data. RESULTS Using high-throughput imaging, more than 300,000 de-identified images of blood films were collected, resulting in more than 1.5 trillion segmented cells. The de-identified images and processed data will be released in an open-access data repository. The frequency distribution of 11 different morphological parameters were used for machine learning based classification, resulting in testing sensitivity and specificity of 82.9% and 82.4%, respectively. The confusion matrix (Figure 1) shows that sensitivity of detecting SCD cases (HbS/β-thalassemia and HbSS) was 98.6%, and most misclassifications were between trait and normal conditions. Comparing the different low-cost techniques (Table 1), the Gazelle Hb variant test had the highest sensitivity and specificity (97.0% and 99.3%), but is more than twice as expensive than the augmented sickling test, while lateral flow assays (HemoTypeSC and SickleSCAN) had lower sensitivity (74-75%) than the augmented sickling test, due to their inability to detect β-thalassemia trait. CONCLUSIONS The sickling test, which is traditionally unable to distinguish between SCT and SCD, was augmented to detect SCD and trait conditions (including β-thalassemia) with an overall sensitivity of 82.4% (98.6% sensitivity for SCD detection). The preliminary results indicate that image-based classification can be a promising low-cost and automated screening tool for low-resource settings, reducing the need for highly-trained personnel for device operation and disease detection. Furthermore, the open-access image dataset can be used to improve classification in the future, as segmentation/classification algorithms improve over time.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,001 | 0,003 |
| Méta-épidémiologie (sens strict) | 0,000 | 0,000 |
| Méta-épidémiologie (sens large) | 0,000 | 0,000 |
| Bibliométrie | 0,001 | 0,001 |
| Études des sciences et des technologies | 0,001 | 0,000 |
| Communication savante | 0,001 | 0,000 |
| Science ouverte | 0,001 | 0,001 |
| Intégrité de la recherche | 0,001 | 0,000 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,002 | 0,000 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».