Empirical evaluation of the Q-Genie tool: a protocol for assessment of effectiveness
Bibliographic record
Abstract
INTRODUCTION: Meta-analyses of genetic association studies are affected by biases and quality shortcomings of the individual studies. We previously developed and validated a risk of bias tool for use in systematic reviews of genetic association studies. The present study describes a larger empirical evaluation of the Q-Genie tool. METHODS AND ANALYSIS: MEDLINE, Embase, Global Health and the Human Genome Epidemiology Network will be searched for published meta-analyses of genetic association studies. Twelve reviewers in pairs will apply the Q-Genie tool to all studies in included meta-analyses. The Q-Genie will then be evaluated on its ability to (i) increase precision after exclusion of low quality studies, (ii) decrease heterogeneity after exclusion of low quality studies and (iii) good agreement with experts on quality rating by Q-Genie. A qualitative assessment of the tool will also be conducted using structured questionnaires. DISCUSSION: This systematic review will quantitatively and qualitatively assess the Q-Genie's ability to identify poor quality genetic association studies. This information will inform the selection of studies for inclusion in meta-analyses, conduct sensitivity analyses and perform metaregression. Results of this study will strengthen our confidence in estimates of the effect of a gene on an outcome from meta-analyses, ultimately bringing us closer to deliver on the promise of personalised medicine. ETHICS AND DISSEMINATION: An updated Q-Genie tool will be made available from the Population Genomics Program website and the results will be submitted for a peer-reviewed publication.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".