A Comparison Among 5 Methods for the Clinical Diagnosis of Fetal Alcohol Spectrum Disorders
Bibliographic record
Abstract
BACKGROUND: Despite the prevalence of fetal alcohol spectrum disorders (FASD) and the importance of accurate identification of patients, clinical diagnosis may not be consistent across sites due to the heterogeneous nature of FASD and the characteristics of different diagnostic systems used. Here, we compare 5 systems designed to operationalize criteria recommended for the diagnosis of effects of prenatal alcohol exposure (PAE). We determined the extent of consistency among them as well as factors that may reduce intersystem reliability. Compared are: Emory Clinic, Seattle 4-Digit System (Diagnostic Guidelines for Fetal Alcohol Spectrum Disorders: The 4-Digit Diagnostic Code, Seattle, WA, University Publication Services, 2004), Centers for Disease Control and Prevention (Fetal Alcohol Syndrome: Guidelines for Referral and Diagnosis, Department of Health and Human Services, Centers for Disease Control and Prevention, Atlanta, GA, 2004), Canadian Guidelines (CMAJ, 172, 2005, S1), and the Hoyme Modifications (Pediatrics, 115, 2005, 39). METHODS: Subjects were 1,581 consecutively registered patients applying for evaluation at a university-based clinic treating alcohol and drug-exposed children. Records of the multidisciplinary evaluation (pediatric, social, psychological) were abstracted. Diagnostic criteria for all 5 systems were applied, and patients were diagnosed according to each of the systems. We compared results using Cohen's Kappa to evaluate the extent of agreement. RESULTS: Percent of individuals diagnosed with FASD ranged from 4.74% (CDC) to 59.58% (Hoyme). Examination using Cohen's Kappa found modest agreement among systems, particularly when individual diagnoses, Fetal Alcohol Syndrome (FAS), partial FAS (pFAS), and Alcohol-Related Neurodevelopmental Disorder (ARND) were used. Examination of diagnostic criteria found almost perfect agreement on growth (weight; height), with limited overlap for physical features (palpebral fissures, hypoplastic philtrum, upper vermillion) and for neurobehavioral outcomes. Child's race and age influenced agreement among systems, with African American and older children more frequently diagnosed. CONCLUSIONS: Results suggest problems in convergent validity among these systems, as demonstrated by a lack of reliability in diagnosis. Absence of an external standard makes it impossible to determine whether any system is more accurate, but outcomes do suggest areas for future research that may refine diagnosis.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.101 | 0.200 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.003 |
| Bibliometrics | 0.009 | 0.004 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.002 | 0.002 |
| Open science | 0.002 | 0.004 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".