Chromosomal microarray analysis of 410 Han Chinese patients with autism spectrum disorder or unexplained intellectual disability and developmental delay
Bibliographic record
Abstract
Abstract Copy number variants (CNVs) are recognized as a crucial genetic cause of neurodevelopmental disorders (NDDs). Chromosomal microarray analysis (CMA), the first-tier diagnostic test for individuals with NDDs, has been utilized to detect CNVs in clinical practice, but most reports are still from populations of European ancestry. To contribute more worldwide clinical genomics data, we investigated the genetic etiology of 410 Han Chinese patients with NDDs (151 with autism and 259 with unexplained intellectual disability (ID) and developmental delay (DD)) using CMA (Affymetrix) after G-banding karyotyping. Among all the NDD patients, 109 (26.6%) carried clinically relevant CNVs or uniparental disomies (UPDs), and 8 (2.0%) had aneuploidies (6 with trisomy 21 syndrome, 1 with 47,XXY, 1 with 47,XYY). In total, we found 129 clinically relevant CNVs and UPDs, including 32 CNVs in 30 ASD patients, and 92 CNVs and 5 UPDs in 79 ID/DD cases. When excluding the eight patients with aneuploidies, the diagnostic yield of pathogenic and likely pathogenic CNVs and UPDs was 20.9% for all NDDs (84/402), 3.3% in ASD (5/151), and 31.5% in ID/DD (79/251). When aneuploidies were included, the diagnostic yield increased to 22.4% for all NDDs (92/410), and 33.6% for ID/DD (87/259). We identified a de novo CNV in 14.9% (60/402) of subjects with NDDs. Interestingly, a higher diagnostic yield was observed in females (31.3%, 40/128) compared to males (16.1%, 44/274) for all NDDs (P = 4.8 × 10−4), suggesting that a female protective mechanism exists for deleterious CNVs and UPDs.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".