Systematic survey of the causal language use in systematic reviews of observational studies: a study protocol
Bibliographic record
Abstract
INTRODUCTION: Sometimes, observational studies may provide important evidence that allow inferences of causality between exposure and outcome (although on most occasions only low certainty evidence). Authors, frequently and perhaps usually at the behest of the journals to which they are submitting, avoid using causal language when addressing evidence from observational studies. This is true even when the issue of interest is the causal effect of an intervention or exposure. Clarity of thinking and appropriateness of inferences may be enhanced through the use of language that reflects the issue under consideration. The objectives of this study are to systematically evaluate the extent and nature of causal language use in systematic reviews of observational studies and to relate that to the actual intent of the investigation. METHODS AND ANALYSIS: We will conduct a systematic survey of systematic reviews of observational studies addressing modifiable exposures and their possible impact on patient-important outcomes. We will randomly select 200 reviews published in 2019, stratified in a 1:1 ratio by use and non-use of the Grading of Recommendations Assessment, Development, and Evaluation (GRADE). Teams of two reviewers will independently assess study eligibility and extract data using a standardised data extraction forms, with resolution of disagreement by discussion and, if necessary, by third party adjudication. Through examining the inferences, they make in their papers' discussion, we will evaluate whether the authors' intent was to address causation or association. We will summarise the use of causal language in the study title, abstract, study question and results using descriptive statistics. Finally, we will assess whether the language used is consistent with the intention of the authors. We will determine whether results in reviews that did or did not use GRADE differ. ETHICS AND DISSEMINATION: Ethics approval for this study is not required. We will disseminate the results through publication in a peer-reviewed journals. REGISTRATION: Open Science Framework (osf.io/vh8yx).
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.403 | 0.512 |
| Meta-epidemiology (narrow) | 0.006 | 0.008 |
| Meta-epidemiology (broad) | 0.013 | 0.014 |
| Bibliometrics | 0.023 | 0.028 |
| Science and technology studies | 0.005 | 0.010 |
| Scholarly communication | 0.008 | 0.015 |
| Open science | 0.005 | 0.009 |
| Research integrity | 0.012 | 0.009 |
| Insufficient payload (model declined to judge) | 0.035 | 0.017 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".