Comprehensive genotype-phenotype analysis in POLR3-related disorders
Bibliographic record
Abstract
RNA polymerase III (RNA Pol III)-related disorders (POLR3-RDs) are a group of clinical entities characterized by causal variants in genes encoding RNA Pol III subunits, including POLR3A, POLR3B, POLR1C, POLR1D, POLR3D, POLR3E, POLR3F, POLR3GL, POLR3H, and POLR3K. These typically cause developmental phenotypes affecting the central nervous system; the eyes; connective tissues including bones, teeth, and endocrine axes; and the reproductive system. Similar phenotypes can be caused by variants in separate subunit genes (multigenic). In contrast, variants in the same gene can cause different phenotypes (pleiotropy), making genotype-phenotype correlation challenging. POLR3-RDs, though individually rare, have never been analyzed collectively. To bridge this gap, we developed an extensive database encompassing all published and unpublished cases of POLR3-RDs and conducted the first comprehensive genotype-phenotype correlation study across their entire spectrum. This work contributed new cases, representing 13% of all documented cases in the literature, along with 31 novel variants, accounting for 8% of all identified variants. This database was constructed by systematically reviewing the literature and integrating data from patients under the care of our international network of collaborators. The dataset includes genotype curation, bioinformatics, prior publications, and individual patient outcome information. By leveraging these comprehensive data, we were able to establish clear genotype-phenotype correlations for some pathogenic variants, which will help provide optimal clinical care and genetic counseling (including insights into disease phenotypes and progression) and offer valuable guidance for future clinical trial design and patient stratification.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".