Ancestry and frequency of genetic variants in the general population are confounders in the characterization of germline variants linked to cancer
Bibliographic record
Abstract
BACKGROUND: Pediatric high-grade gliomas (pHGGs) are incurable malignant brain cancers. Clear somatic genetic drivers are difficult to identify in the majority of cases. We hypothesized that this may be due to the existence of germline variants that influence tumor etiology and/or progression and are filtered out using traditional pipelines for somatic mutation calling. METHODS: In this study, we analyzed whole-genome sequencing (WGS) datasets of matched germlines and tumor tissues to identify recurrent germline variants in pHGG patients. RESULTS: We identified two structural variants that were highly recurrent in a discovery cohort of 8 pHGG patients. One was a ~ 40 kb deletion immediately upstream of the NEGR1 locus and predicted to remove the promoter region of this gene. This copy number variant (CNV) was present in all patients in our discovery cohort (n = 8) and in 86.3% of patients in our validation cohort (n = 73 cases). We also identified a second recurrent deletion 55.7 kb in size affecting the BTNL3 and BTNL8 loci. This BTNL3-8 deletion was observed in 62.5% patients in our discovery cohort, and in 17.8% of the patients in the validation cohort. Our single-cell RNA sequencing (scRNA-seq) data showed that both deletions result in disruption of transcription of the affected genes. However, analysis of genomic information from multiple non-cancer cohorts showed that both the NEGR1 promoter deletion and the BTNL3-8 deletion were CNVs occurring at high frequencies in the general population. Intriguingly, the upstream NEGR1 CNV deletion was homozygous in ~ 40% of individuals in the non-cancer population. This finding was immediately relevant because the affected genes have important physiological functions, and our analyses showed that NEGR1 expression levels have prognostic value for pHGG patient survival. We also found that these deletions occurred at different frequencies among different ethnic groups. CONCLUSIONS: Our study highlights the need to integrate cancer genomic analyses and genomic data from large control populations. Failure to do so may lead to spurious association of genes with cancer etiology. Importantly, our results showcase the need for careful evaluation of differences in the frequency of genetic variants among different ethnic groups.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".