MétaCan
Menu
Back to cohort
Record W4403437675 · doi:10.2196/64995

The Prevalence of Sickle Cell Disease in Colorado and Methodologies of the Colorado Sickle Cell Data Collection Program: Public Health Surveillance Study

2024· article· en· W4403437675 on OpenAlexvenueno aff
Joshua I Miller, Kathryn Hassell, Yvonne Kellar‐Guenther, Stacey Quesada, Rhonda West, Marci K. Sontag

Bibliographic record

VenueJMIR Public Health and Surveillance · 2024
Typearticle
Languageen
FieldMedicine
TopicHemoglobinopathies and Related Disorders
Canadian institutionsnot available
Fundersnot available
KeywordsPublic healthPreprintEnvironmental healthMedicinePublic health surveillanceDiseaseData collectionComputer sciencePathologyWorld Wide WebStatistics

Abstract

fetched live from OpenAlex

Background: Sickle cell disease (SCD) is a genetic blood disorder that affects approximately 100,000 individuals in the United States, with the highest prevalence among Black or African American populations. While advances in care have improved survival, comprehensive state-level data on the prevalence of SCD remain limited, which hampers efforts to optimize health care services. To address this gap, the Colorado Sickle Cell Data Collection (CO-SCDC) program was established in 2021 as part of the Centers for Disease Control and Prevention's initiative to enhance surveillance and public health efforts for SCD. Objective: The objectives of this study were to describe the establishment of the CO-SCDC program and to provide updated estimates of the prevalence and birth prevalence of SCD in Colorado, including geographic dispersion. Additional objectives include evaluating the accuracy of case identification methods and leveraging surveillance activities to inform public health initiatives. Methods: Data were collected from Health Data Compass (a multi-institutional data warehouse) containing electronic health records from the University of Colorado Health and Children's Hospital Colorado for the years 2012-2020. Colorado newborn screening program data were included for confirmed SCD diagnoses from 2001 to 2020. Records were linked using the Colorado University Record Linkage tool and deidentified for analysis. Case definitions, adapted from the Centers for Disease Control and Prevention's Registry and Surveillance System for Hemoglobinopathies project, classified cases as possible, probable, or definite SCD. Clinical validation by hematologists was performed to ensure accuracy, and prevalence rates were calculated using 2020 US Census population estimates. Results: In 2019, 435 individuals were identified as living with SCD in Colorado, an increase of 16%-40% over previous estimates, with the majority (n=349, 80.2%) identifying as Black or African American. The median age of individuals was 19 years. The prevalence of SCD was highest in urban counties, with concentrations in Arapahoe, Denver, and El Paso counties. Birth prevalence of SCD increased from 11.9 per 100,000 live births between 2010 and 2014 to 20.1 per 100,000 live births between 2015 and 2019 with 58.5% (n=38) of cases being hemoglobin (Hb) SS or HbSβ0 thalassemia subtypes. The study highlighted a 67% (n=26) increase in SCD births over the decade, correlating with the growth of the Black or African American population in the state. Conclusions: The CO-SCDC program successfully established the capacity to perform SCD surveillance and, in doing so, identified baseline prevalence estimates for SCD in Colorado. The findings highlight geographic dispersion across Colorado counties, highlighting the need for equitable access to specialty care, particularly for rural populations. The combination of automated data linkage and clinical validation improved case identification accuracy. Future efforts will expand surveillance to include claims data to better capture health care use and address potential underreporting. These results will guide public health interventions aimed at improving care for individuals with SCD in Colorado.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.002
metaresearch head score (Gemma)0.005
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: Observational
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.396
Threshold uncertainty score0.788

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0020.005
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0000.000
Bibliometrics0.0020.003
Science and technology studies0.0010.000
Scholarly communication0.0010.001
Open science0.0010.001
Research integrity0.0000.001
Insufficient payload (model declined to judge)0.0010.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.070
GPT teacher head0.362
Teacher spread0.291 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designObservational
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations2
Published2024
Admission routes1
Has abstractyes

Explore more

Same venueJMIR Public Health and SurveillanceSame topicHemoglobinopathies and Related DisordersFrench-language works237,207