OC23.05: Prenatal diagnosis of fetal musculoskeletal disorders: how accurate can we be?
Bibliographic record
Abstract
To determine the accuracy of prenatal ultrasound and whether knowledge of specific pathological findings can further improve accuracy. To determine clinical applicability of Nosology 2010 classification system. A retrospective review of perinatal autopsies from 2002–2011 was performed to identify cases with a diagnosis of a musculoskeletal disorder, as defined by Nosology 2010. An initial read of the corresponding images and reports were performed. A repeat evaluation was performed after the reviewer was given access to the pathology findings, to determine if diagnostic accuracy could be improved by knowledge of the specific pathological findings. Data was entered into a standardised web-based reporting system. Accuracy of the initial ultrasound diagnosis was compared to the pathology diagnosis. Findings of the repeat evaluation were cross indexed to the initial read, to determine which findings may be missed. A musculoskeletal disorder was identified in 112 of 2002 (5.6%) perinatal autopsies. 91 cases had both pathology and ultrasound imaging available. The 91 cases encompassed 16 of the 40 Nosology 2010 groups. The most common specific diagnoses were thanatophoric dysplasia type 1 or 2 at 20 (22%) and osteogenesis imperfecta type 2 at 18 (20%). Accurate assessment of lethality was achieved in 90/91 (99%) of cases on both the initial and the re-read ultrasound. An accurate group diagnosis was achieved in 66/91 (73%) on the initial ultrasound, with an increased to 76/91 (84%) on the re-read ultrasound. An accurate sub-group diagnosis was achieved in 48/91 (53%) on initial read and 55/91 (60%) on the re-read. Prenatal ultrasound is extremely accurate for the diagnosis of a lethal musculoskeletal disorder, can achieve a correct group diagnosis in 84% and a correct sub-group diagnosis in 60%. Prior knowledge of the pathology results helps identify features that may have been missed on the initial ultrasound. This may lead to improved prospective evaluation in these rare disorders.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.028 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".