Reanalysing genomic data by normalized coverage values uncovers CNVs in bone marrow failure gene panels
Bibliographic record
Abstract
Abstract Inherited bone marrow failure syndromes (IBMFSs) are genetically heterogeneous disorders with cytopenia. Many IBMFSs also feature physical malformations and an increased risk of cancer. Point mutations can be identified in about half of patients. Copy number variation (CNVs) have been reported; however, the frequency and spectrum of CNVs are unknown. Unfortunately, current genome-wide methods have major limitations since they may miss small CNVs or may have low sensitivity due to low read depths. Herein, we aimed to determine whether reanalysis of NGS panel data by normalized coverage value could identify CNVs and characterize them. To address this aim, DNA from IBMFS patients was analyzed by a NGS panel assay of known IBMFS genes. After analysis for point mutations, heterozygous and homozygous CNVs were searched by normalized read coverage ratios and specific thresholds. Of the 258 tested patients, 91 were found to have pathogenic point variants. NGS sample data from 165 patients without pathogenic point mutations were re-analyzed for CNVs; 10 patients were found to have deletions. Diamond Blackfan anemia genes most commonly exhibited heterozygous deletions, and included RPS19 , RPL11 , and RPL5 . A diagnosis of GATA2 -related disorder was made in a patient with myelodysplastic syndrome who was found to have a heterozygous GATA2 deletion. Importantly, homozygous FANCA deletion were detected in a patient who could not be previously assigned a specific syndromic diagnosis. Lastly, we identified compound heterozygousity for deletions and pathogenic point variants in RBM8A and PARN genes. All deletions were validated by orthogonal methods. We conclude that careful analysis of normalized coverage values can detect CNVs in NGS panels and should be considered as a standard practice prior to do further investigations.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".