MétaCan
Menu
Back to cohort
Record W4389243360 · doi:10.1182/blood-2023-187754

Low-Cost Automated Microscopy and Morphology-Based Machine Learning Classification of Sickle Cell Disease and Beta-Thalassemia in Nepal and Canada

2023· article· en· W4389243360 on OpenAlexaffabout
Pranav Shrestha, Hendrik Lohse, Christopher Bhatla, Rodrigo Onell, Nicholas Au, Ali Amid, Heather McCartney, Alaa Alzaki, Mykola Maydan, Navdeep Sandhu, Hayley Merkeley, Videsh Kapoor, Rajan Pande, Boris Stoeber

Bibliographic record

VenueBlood · 2023
Typearticle
Languageen
FieldMedicine
TopicHemoglobinopathies and Related Disorders
Canadian institutionsBC Children's HospitalSt. Paul's HospitalUniversity of British Columbia
Fundersnot available
KeywordsMedicineSickle cell traitThalassemiaPediatricsMalariaAlpha-thalassemiaAsymptomaticGold standard (test)DiseaseInternal medicinePathologyGenotypeBiologyGenetics

Abstract

fetched live from OpenAlex

INTRODUCTION Sickle cell disease (SCD) is the most common inherited blood disorder, and is associated with a high mortality rate for children in low-resource settings, where the disease burden is high. Early screening and treatment options reduce morbidity and mortality. In addition to screening for SCD, screening for asymptomatic, heterozygous conditions, such as sickle cell trait (SCT) and β-thalassemia trait, is required for effective SCD and thalassemia screening programs. However, rural and remote communities lack accurate low-cost screening techniques. The goal of this study is to augment a common low-cost screening test, the sickling test, using automated microscopy and machine learning to classify individuals with SCD (HbS/β-thalassemia and HbSS), trait conditions (HbAS and HbA/β-thalassemia), and those without common beta-globin disorders (HbAA) for implementation in Nepal. As HbC mutation is not encountered in Nepal, screening for HbSC or HbAC was not included. Additionally, we aim to evaluate the feasibility of 5 low-cost techniques to supplement/replace the local gold standard test (Hb HPLC), which is inaccessible in rural/remote settings. METHODS In this study (ClinicalTrials.gov Identifier: NCT05506358), blood samples were collected and tested at respective clinical sites in Nepal (Mount Sagarmatha Polyclinic and Diagnostic Center, Nepalgunj) and Canada (St. Paul's Hospital and BC Children's Hospital, Vancouver). Participants (138 total: 111 Nepal; 27 Canada) between ages 2 to 74 whose diagnoses were previously established by Hb HPLC were recruited. Five low-cost tests (our augmented sickling test, HbS solubility test, HemoTypeSC, Sickle SCAN, Gazelle Hb variant test) and Hb HPLC were performed on all participants (30 HbAA; 23 HbA/β-thalassemia; 45 HbAS; 11 HbS/β-thalassemia; 29 HbSS). Informed consent was provided by the participants or parents, according to protocols approved by institutional/national research ethics boards. For the sickling test, a sealed wet preparation of blood mixed with 2% sodium metabisulphite was imaged with a robotic microscope (Octopi) with automated scanning/focusing. The images of blood cells were segmented using Cellpose 2.0, allowing for morphological characterization of cells and morphology based classification, with an 80:20 participant-wise split of training and testing data. RESULTS Using high-throughput imaging, more than 300,000 de-identified images of blood films were collected, resulting in more than 1.5 trillion segmented cells. The de-identified images and processed data will be released in an open-access data repository. The frequency distribution of 11 different morphological parameters were used for machine learning based classification, resulting in testing sensitivity and specificity of 82.9% and 82.4%, respectively. The confusion matrix (Figure 1) shows that sensitivity of detecting SCD cases (HbS/β-thalassemia and HbSS) was 98.6%, and most misclassifications were between trait and normal conditions. Comparing the different low-cost techniques (Table 1), the Gazelle Hb variant test had the highest sensitivity and specificity (97.0% and 99.3%), but is more than twice as expensive than the augmented sickling test, while lateral flow assays (HemoTypeSC and SickleSCAN) had lower sensitivity (74-75%) than the augmented sickling test, due to their inability to detect β-thalassemia trait. CONCLUSIONS The sickling test, which is traditionally unable to distinguish between SCT and SCD, was augmented to detect SCD and trait conditions (including β-thalassemia) with an overall sensitivity of 82.4% (98.6% sensitivity for SCD detection). The preliminary results indicate that image-based classification can be a promising low-cost and automated screening tool for low-resource settings, reducing the need for highly-trained personnel for device operation and disease detection. Furthermore, the open-access image dataset can be used to improve classification in the future, as segmentation/classification algorithms improve over time.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.001
metaresearch head score (Gemma)0.003
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Bench or experimental · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.635
Threshold uncertainty score0.725

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0010.003
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0000.000
Bibliometrics0.0010.001
Science and technology studies0.0010.000
Scholarly communication0.0010.000
Open science0.0010.001
Research integrity0.0010.000
Insufficient payload (model declined to judge)0.0020.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.007
GPT teacher head0.232
Teacher spread0.225 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designBench or experimental
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations3
Published2023
Admission routes2
Has abstractyes

Explore more

Same venueBloodSame topicHemoglobinopathies and Related DisordersFrench-language works237,207