Utility and limitations of genetic disease databases in clinical genetics research: A neurofibromatosis 1 database example
Bibliographic record
Abstract
Databases that collect clinical information on patients with particular genetic diseases can be used to investigate the clinical history of a disorder, its genetics, and genotype-phenotype correlations. A database can also serve as a valuable source of patients for studies of disease pathogenesis, variability, or treatment. We review the strengths and limitations of genetic disease databases in the context of our experience with the National Neurofibromatosis Foundation International Database (NNFFID). Genetic disease databases have been developed by individual investigators, scientific consortia, patient support organizations, and commercial enterprises. Databases vary from simple lists of affected individuals to comprehensive collections of detailed clinical and genetic information. Data may be obtained from people who volunteer to be included, systematic assessments of patients seen at participating medical centers, or population-based registries. Access to information may be highly restricted or widely available. These variables all affect the possible uses and usefulness of the data for research. Technical aspects of data entry, organization, storage, and retrieval, as well as issues related to data quality, confidentiality, and security, help determine how well a system actually functions. We discuss examples of research that have been accomplished with genetic disease databases and make recommendations regarding the organization and operation of these resources.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.006 | 0.018 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.002 | 0.000 |
| Bibliometrics | 0.001 | 0.002 |
| Science and technology studies | 0.000 | 0.005 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.002 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".