Impact of Protocol Access on Bias Ratings in Cohort Studies of Interventions: A Before-and-After Study
Bibliographic record
Abstract
This single-arm before-and-after study aims to assess estimate how often ROBINS-I risk-of-bias ratings based on published study reports are changed following access to study protocols. Secondary aim is to estimate whether access to study protocols is associated with improved inter-rater reliability in ROBINS-I domain-level bias ratings. Study materials involved are protocols and publications of cohort studies of interventions, defined as controlled retrospective or prospective cohort studies that assess the effects of interventions (e.g., drugs, biologics, devices, procedures, policies, counselling) that aim to affect health-related outcomes. They also need to have been registered on ClinicalTrials.gov with a prespecified study protocol (protocol version date is within one month after the listed study start date), have a study status marked as completed as well as a study start date between 2016 and 2020, and have a publication of results in a peer-reviewed journal. Eligible assessors, namely the people who will be rating the bias, are the authors of published Cochrane systematic reviews that applied the ROBINS-I to cohort studies of interventions. If additional assessors are needed, secondary sources will include authors of other published systematic reviews using ROBINS-I for cohort studies, and staff members from the systematic review teams at Cochrane Denmark and Canada. Included studies will be randomly assigned to assessors using a balanced incomplete block (BIB) design, such that each assessor rates the same number of studies, each study is rated by the same number of assessors, and each distinct pair of assessors would have the same number of studies in common. This design substantially reduces rating burden by having each assessor review only a subset of studies, while still maintaining estimability of inter-reader agreement for all pairs of assessors. In each study-assessor pair (e.g., Assessor A’s rating of Study #1), we will identify the differences in domain-level bias ratings made before and after protocol access.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.019 | 0.022 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.004 | 0.001 |
| Bibliometrics | 0.005 | 0.015 |
| Science and technology studies | 0.000 | 0.007 |
| Scholarly communication | 0.002 | 0.002 |
| Open science | 0.011 | 0.014 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.005 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; both teacher heads agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".