Assessment of the functionality and usability of open-source rare variant analysis pipelines
Bibliographic record
Abstract
Sequencing of increasingly larger cohorts has revealed many rare variants, presenting an opportunity to further unravel the genetic basis of complex traits. Compared with common variants, rare variants are more complex to analyze. Specialized computational tools for these analyses should be both flexible and user-friendly. However, an overview of the available rare variant analysis pipelines and their functionalities is currently lacking. Here, we provide a systematic review of the currently available rare variant analysis pipelines. We searched MEDLINE and Google Scholar until 27 November 2023, and included open-source rare variant pipelines that accepted genotype data from cohort and case-control studies and group variants into testing units. Eligible pipelines were assessed based on functionality and usability criteria. We identified 17 rare variant pipelines that collectively support various trait types, association tests, testing units, and variant weighting schemes. Currently, no single pipeline can handle all data types in a scalable and flexible manner. We recommend different tools to meet diverse analysis needs. STAARpipeline is suitable for newcomers and common applications owing to its built-in definitions for the testing units. REGENIE is highly scalable, actively maintained, regularly updated, and well documented. Ravages is suitable for analyzing multinomial variables, and OrdinalGWAS is tailored for analyzing ordinal variables. Opportunities remain for developing a user-friendly pipeline that provides high degrees of flexibility and scalability. Such a pipeline would enable researchers to exploit the potential of rare variant analyses to uncover the genetic basis of complex traits.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".