Clinical significance of preleukemic somatic <i>GATA1</i> mutations in children with Down syndrome
Bibliographic record
Abstract
ABSTRACT: Children with Down syndrome (DS) have a high risk of GATA1-associated myeloid leukemia (ML-DS) before age 4 years. Somatic N-terminal GATA1 mutations (GATA1s) are necessary, but not sufficient, for ML-DS, but their significance at birth for individual babies and whether mutations occur after birth is unclear. To address these questions, we performed a prospective study of newborns with DS using next-generation sequencing-based GATA1 mutation analysis, with hematologic and clinical evaluation and follow-up for the window of ML-DS risk. Of 450 neonates with DS, 113 (25%) had GATA1s mutations, among whom 20/113 (17.7%) had multiple mutations and 59 (52%) were clinically silent. Variant allele frequency (VAF) varied from 0.3% to 89%. VAF positively correlated (P < .0001) with the percent blasts, leukocytes, dyserythropoiesis and dysmegakaryopoiesis scores, and clinical disease severity, and negatively with hemoglobin, although only 4/113 were anemic. GATA1s mutations were detected from 28 weeks gestation; the highest frequency (45%) was at 34 to 35 weeks, whereas mutation frequency in early fetal samples (<20 weeks) was <4% (2/57). GATA1s clones (VAF, percent blasts) fell rapidly postnatally, becoming undetectable by 6 months, except in neonates who developed ML-DS. Of 110 surviving neonates, 7 (6.4%) developed ML-DS at a median age of 17.5 months. GATA1s clone size at birth was the only predictor of ML-DS. No neonates lacking GATA1s mutations acquired mutations after birth or developed ML-DS. Taken together, the fetal environment is essential for GATA1s mutation selection and expansion of GATA1s clones. Rates of leukemic transformation of GATA1s clones detected at birth are low, but clones that persist >6 months transformed.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".