Maximizing Benefits and Minimizing Harms: Diagnostic Uncertainty Arising From Newborn Screening
Bibliographic record
Abstract
Screening newborns for rare diseases for which effective treatments are available is a crucial public health practice. Indeed, the expansion of population-wide newborn screening (NBS) programs was named by the Centers for Disease Control and Prevention as 1 of the 10 great public health achievements in the first 10 years of the 21st century.1 Pivotal randomized controlled trials of NBS for cystic fibrosis (CF) revealed potential benefits,2,3 with evidence from the Wisconsin trial revealing that early ascertainment through screening could prevent severe malnutrition and improve long-term growth.2 NBS for CF is now widely implemented in many jurisdictions and is included in the influential US Recommended Uniform Screening Panel.4As Gray et al5 provocatively noted, however, “All screening programmes do harm; some do good as well, and, of these, some do more good than harm at reasonable cost. The first task of any public health service is to identify beneficial programmes by appraising the evidence.”5(p480) The current study by Gonska et al6 is an important example of ongoing research required to measure benefits and understand harms of NBS, thus supporting program improvement.Although the benefit of NBS for CF is unequivocal,7 there are potential harms to children and their families who receive false-positive or inconclusive results after a positive screen result. Although most false-positive findings are resolved in early infancy, with evidence mainly revealing transient psychosocial impacts that are mitigated with timely follow-up and effective communication,8,9 inconclusive results tend to be longer lasting. This is indeed the case for children identified through NBS as having a diagnosis of
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.191 | 0.509 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.004 | 0.002 |
| Bibliometrics | 0.004 | 0.003 |
| Science and technology studies | 0.006 | 0.028 |
| Scholarly communication | 0.014 | 0.020 |
| Open science | 0.005 | 0.012 |
| Research integrity | 0.013 | 0.017 |
| Insufficient payload (model declined to judge) | 0.005 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".