The LORIS MyeliNeuroGene rare disease database for natural history studies and clinical trial readiness
Bibliographic record
Abstract
BACKGROUND: Rare diseases are estimated to affect 150-350 million people worldwide. With advances in next generation sequencing, the number of known disease-causing genes has increased significantly, opening the door for therapy development. Rare disease research has therefore pivoted from gene discovery to the exploration of potential therapies. With impending clinical trials on the horizon, researchers are in urgent need of natural history studies to help them identify surrogate markers, validate outcome measures, define historical control patients, and design therapeutic trials. RESULTS: We customized a browser-accessible multi-modal (e.g. genetics, imaging, behavioral, patient-determined outcomes) database to increase cohort sizes, identify surrogate markers, and foster international collaborations. Ninety data entry forms were developed including family, perinatal, developmental history, clinical examinations, diagnostic investigations, neurological evaluations (i.e. spasticity, dystonia, ataxia, etc.), disability measures, parental stress, and quality of life. A customizable clinical letter generator was created to assist in continuity of patient care. CONCLUSIONS: Small cohorts and underpowered studies are a major challenge for rare disease research. This online, rare disease database will be accessible from all over the world, making it easier to share and disseminate data. We have outlined the methodology to become Title 21 Code of Federal Regulations Part 11 Compliant, which is a requirement to use electronic records as historical controls in clinical trials in the United States. Food and Drug Administration compliant databases will be life-changing for patients and families when historical control data is used for emerging clinical trials. Future work will leverage these tools to delineate the natural history of several rare diseases and we are confident that this database will be used on a larger scale to improve care for patients affected with rare diseases.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.003 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".