Comparison of three systems for the diagnosis of fetal alcohol spectrum disorders in a community sample
Bibliographic record
Abstract
BACKGROUND: It is estimated that 1%-5% of children in the United States are affected by prenatal alcohol exposure while only a small percentage are so identified in clinical practice. One explanation for this discrepancy may be the way in which diagnostic criteria are operationalized. METHODS: To evaluate the extent to which three commonly used systems for the diagnosis of Fetal Alcohol Spectrum Disorder (FASD) consistently identified children in a community sample, data from the Collaboration on Fetal Alcohol Spectrum Disorders Prevalence (COFASP) study were re-analyzed. In the data set, there were 2325 children with variables necessary to allow diagnosis by three systems commonly used in North America. These systems were (1) that used by COFASP, which is a revised modification of the Institute of Medicine's recommendations, (2) the 4-Digit Code, and (3) the most recent Canadian Guidelines. To determine the degree of association among these classifications, the Fleiss Multirater Kappa measure of agreement was applied. RESULTS: Among these three systems, 408 children were classified as FASD, 208 by the CoFASP system, 319 by the 4-Digit Code, and 28 by the Canadian Guidelines. Agreement among the findings from the three systems varied from slight to fair. CONCLUSIONS: These results indicate a lack of consistency in these approaches to FASD diagnosis. Discrepancies result from differences in specifying the criteria used to define the diagnosis, including growth, physical features, neurobehavior, and alcohol-use thresholds. The question of their relative accuracy cannot be resolved without reference to a measure of validity that does not currently exist, and this suggests the need for a more empirically based diagnostic schema.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.005 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".