Identification and Classification of Rare Variants in NPC1 and NPC2 in Quebec
Bibliographic record
Abstract
Niemann-Pick disease type C (NPC) is a treatable autosomal recessive neurodegenerative condition which leads to a variety of progressive manifestations. Despite most cases being diagnosed at a young age, disease prevalence may be underestimated, especially in adults, and interpretation of NPC1 and NPC2 variants can be difficult. This study aims to identify potential pathogenic variants in a large cohort of healthy individuals and classify their risk of pathogenicity to assist with future interpretation of variants. The CARTaGENE (CaG) cohort was used to identify possible variants of NPC1 and NPC2. Nine-hundred and eleven RNA samples and 198 exome sequencing were screened for genetic variants through a bio-informatic pipeline performing alignment and variant calling. The identified variants were analyzed using annotations for allelic frequency, pathogenicity and conservation scores. The ACMG guidelines were used to classify the variants. These were then compared to existing databases and previous studies of NPC prevalence, including the Tübingen NPC database. Thirty-two distinct variants were identified after running the samples in the RNA-sequencing pipeline, two of which were classified as pathogenic and 21 of which were not published previously. Furthermore, 46 variants were both identified in our population and with the Tübingen database, the majority of which were of uncertain significance. Ten additional variants were found in our exome-sequencing sample. This study of a sample from a population living in Quebec demonstrates a variety of rare variants, some of which were already described in the literature as well as some novel variants. Classifying these variants is arduous given the scarcity of available literature, even so in a population of healthy individuals. Yet using this data, we were able to identify two pathogenic variants within our population and several new variants not previously identified.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.003 | 0.002 |
| Science and technology studies | 0.002 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.005 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".