Application of observational research methods to real-world studies for rare disease drugs: A scoping review protocol
Bibliographic record
Abstract
The primary objective is to identify which observational research methods have been used in the last 5 years in rare disease drug evaluation and how they are applied to generate adequate evidence regarding the real-world effectiveness or safety of rare disease drugs. Rare disease is an umbrella term for a condition which affects < 200,000 people each year and despite the rarity of these conditions, collectively they encompass approximately 7000 different conditions. With the striking number of rare conditions, many pharmaceutical manufacturers are introducing an increased number of drugs to treat them. However, due to small patient populations, heterogeneity and other factors related to rare diseases, there are feasibility concerns regarding the generation of adequate efficacy and safety evidence using conventional randomized controlled trials (RCTs). Recently, real-world evidence generated through observational (or real-world) studies has been proposed to address some of the feasibility concerns with RCTs by measuring drug effectiveness or safety in the real-world setting. However, there remain methodological concerns due to a lack of randomization/masking. This proposed scoping review aims to identify which observational research methods in the last 5 years are used in rare disease drug evaluation to address methodological concerns and how they are applied to generate evidence on drug effectiveness or safety. Articles must be primary observational or real-world studies reporting rare disease drug effectiveness or safety published within the five years preceding this review. Literature reviews, meta-analyses, randomized control trials, case series, case reports, opinion pieces, conference abstracts, and studies with unavailable full-text articles will be excluded. The search strategy will combine the following key search concepts: rare disease, drugs for rare disease and observational/real-world studies. The search will be conducted in MEDLINE and EMBASE. Review registration number: Open Science Framework, https://osf.io/f3wpv.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.184 | 0.146 |
| Meta-epidemiology (narrow) | 0.005 | 0.006 |
| Meta-epidemiology (broad) | 0.012 | 0.013 |
| Bibliometrics | 0.019 | 0.015 |
| Science and technology studies | 0.005 | 0.006 |
| Scholarly communication | 0.009 | 0.009 |
| Open science | 0.007 | 0.009 |
| Research integrity | 0.011 | 0.009 |
| Insufficient payload (model declined to judge) | 0.071 | 0.020 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".