Accuracy of Expected Risk of Down Syndrome Using the Second-Trimester Triple Test
Bibliographic record
Abstract
Second-trimester maternal serum screening (MSS) for Down syndrome has been widely used in routine prenatal care in developed countries. The screening combines maternal age-specific risk of Down syndrome with risk estimation obtained by measuring maternal serum markers to assign women an expected risk of having a term Down syndrome pregnancy. Diagnostic tests were offered to women whose risk exceeded the risk cutoff determined by the screening program. The commonly used triple test, which involves the use of maternal age, serum α-fetoprotein, unconjugated estriol, and human chorionic gonadotropin, was expected to have a Down syndrome detection rate of 60–65% and false-positive rate of 5% (1). Although the expected screening performance has been achieved in many screening programs, the accuracy of individual risk calculated by a relatively complex computation based on a statistical model was not immediately obvious. Good agreement between the expected risk of Down syndrome and observed prevalence has been reported previously in several screening programs (2)(3)(4)(5). We evaluated the accuracy of expected risk of Down syndrome in a large provincial, multiple test center, MSS program in Ontario, Canada. MSS has been coordinated at the provincial level in Ontario since 1993. Triple maker screening (α-fetoprotein, unconjugated estriol, and β-human chorionic gonadotropin) was carried out in seven regional laboratory centers. Information including screen utilization, results, follow-up data, and the pregnancy outcomes of all women screened in the seven centers was collected …
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.014 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".