Use of Instrumental Variable Analyses for Evaluating Comparative Effectiveness in Empirical Applications of Oncology: A Systematic Review
Bibliographic record
Abstract
PURPOSE: This systematic review aims to characterize the use and trends of instrumental variables (IVs) in oncology research, assess the quality and completeness of IV reporting, and evaluate the agreement and interpretation of IV results in comparison with other techniques used for determining comparative effectiveness in observational research. METHODS: We performed a systematic search of observational empirical oncology papers evaluating the comparative effectiveness of cancer treatments using IV methods. EMBASE and MEDLINE (through June 2021) were used for a keyword search; Scopus and Web of Science were used for a citation search. Publication details and characteristics of IV analysis and reporting were extracted from each study to examine the uptake and quality of IV applications. RESULTS: Sixty-five empirical papers were identified from February 2001 through June 2021. Geographic variation (50.8%) was the most common type of IV used, and the majority of IV applications constructed binary instruments (53.8%). Concurrent analyses using another non-IV method to adjust for confounding were conducted in 56 (86.2%) studies, 17 (30.4%) of which produced results divergent from IV approaches. We observed a modest uptake of IV methods between 2011 and 2021 together with its dissemination, which remained fairly limited to the United States (76.9%). The quality and completeness of IV reporting varied greatly. The underlying assumptions required for a valid IV analysis were only accounted for in full by 20 (30.8%) studies. CONCLUSION: There are limited use and variable quality of IV analyses in oncology. Future research should look to establish standards to better facilitate the quality, transparency, and completeness of IV reporting in this setting.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.292 | 0.663 |
| Meta-epidemiology (narrow) | 0.003 | 0.002 |
| Meta-epidemiology (broad) | 0.017 | 0.027 |
| Bibliometrics | 0.026 | 0.027 |
| Science and technology studies | 0.001 | 0.005 |
| Scholarly communication | 0.009 | 0.008 |
| Open science | 0.005 | 0.004 |
| Research integrity | 0.004 | 0.003 |
| Insufficient payload (model declined to judge) | 0.003 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".