Identifying the ‘incredible’! Part 2: Spot the difference - a rigorous risk of bias assessment can alter the main findings of a systematic review
Bibliographic record
Abstract
Systematic reviews are a valuable tool to inform healthcare decision-making.1 2 While a single randomised controlled trial (RCT) is insufficient to definitively guide healthcare decisions, a systematic review synthesising multiple RCTs can overcome this limitation. The results of rigorous systematic reviews possess wide-ranging applicability to numerous stakeholders within the evidence-based medicine ‘ecosystem’. Clinicians consult systematic reviews to inform their clinical decisions.3 Researchers rely on systematic reviews to identify knowledge gaps in existing literature.4 Health policymakers use systematic review evidence to inform practice guidelines and legislation.5 6 Journal editors often prioritise systematic reviews for their impact on readership attention and journal metrics.7 Finally, patients are empowered by systematic reviews that assess the beneficial and harmful patient-important outcomes of available management strategies.8 Evidently, systematic review authors have an important responsibility to ensure their findings provide the most accurate results possible.The biomedical literature expands by 22 systematic reviews daily,9 with no evidence that production is waning. More systematic reviews are desirable if they identify and inform important research questions that improve patient care.10 However, production of this magnitude is problematic when systematic reviews offer ‘extensive redundancy, little value, misleading claims and/or vested interests’.11 As we outlined in part 1, bias is a systematic deviation from the truth in the results of a research study due to limitations in study design, conduct, or analysis.2 Deviations may either overestimate or underestimate a study’s true findings depending of the type and magnitude of bias. As the results of a systematic review are only as valid as the studies it includes, pooling biased results from different studies can compromise the credibility of systematic review findings when no assessment, or a poor assessment, of risk of bias is performed.3 12Inadequate study design, conduct, or analysis …
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.535 | 0.847 |
| Meta-epidemiology (narrow) | 0.005 | 0.005 |
| Meta-epidemiology (broad) | 0.013 | 0.009 |
| Bibliometrics | 0.013 | 0.008 |
| Science and technology studies | 0.004 | 0.014 |
| Scholarly communication | 0.019 | 0.020 |
| Open science | 0.008 | 0.009 |
| Research integrity | 0.025 | 0.015 |
| Insufficient payload (model declined to judge) | 0.018 | 0.006 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".