Reporting quality of systematic review protocols in environmental and occupational toxicology: a meta-epidemiology study protocol
Bibliographic record
Abstract
Background Systematic reviews (SR), an evidence-synthesis tool, have been increasingly used in toxicology, for regulatory decision-making and allocation of research funding resources. A critical domain of methodological quality is publicly available in the SR protocol that describes the review methods prior to the review. Earlier work showed that only a small number of published SRs in environmental health had identifiable protocols, but did not evaluate their completeness of reporting. Thus, there is a critical need to assess the number of protocols published in peer-reviewed literature related to environmental/occupational toxicology and how much those SR protocols adhere to current reporting standards such as PRISMA-P.Objective The overarching objective of this review is to assess the reporting quality of SR protocols in environmental and occupational toxicology, published in peer-reviewed journals.Methods and analysis This protocol has been reported following PRISMA-P and PRISMA-S checklist and study selection, data extraction, and reporting quality assessment were piloted by independent reviewers. Four electronic databases (PubMed, Scopus, Web of Science, EMBASE) were selected and will be searched for SR protocols using pre-defined strings. Both title/abstract and full-text screening will follow detailed eligibility criteria based on the population concept context (PCC) framework. We will include peer-reviewed, self-identified SR protocols that aim to assess the adverse effects of environmental and occupational exposures. We will perform data extraction using the form that includes general bibliographic information, exposure, adverse effect information, evidence streams, and guidelines/checklists used in protocol reporting or development. A reporting assessment form consisting of 15 PRISMA-P-based items, of which eleven with identical binary responses (reported vs not reported) address critical elements of SR protocol. Assessments of the eleven quality items will provide data for the analysis of the reporting quality of SR protocols. We employ reporting quality as it provides a practical and suitable approach to evaluating the inclusion of essential SR elements, under the assumption that these elements would be reported if they had been appropriately planned rather than overlooked. The screening and data extraction will be conducted by two independent reviewers. Disagreements will be discussed if needed with the support of a third reviewer. The data analysis will describe the total number of SR protocols and trends over time, and absolute frequencies and/or proportion of bibliographic information items. It will summarize the reporting quality of all included protocols by a histogram and rank the quality criteria by their level of adherence across the included protocols.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.458 | 0.668 |
| Meta-epidemiology (narrow) | 0.007 | 0.009 |
| Meta-epidemiology (broad) | 0.018 | 0.021 |
| Bibliometrics | 0.031 | 0.027 |
| Science and technology studies | 0.006 | 0.010 |
| Scholarly communication | 0.014 | 0.014 |
| Open science | 0.007 | 0.010 |
| Research integrity | 0.014 | 0.012 |
| Insufficient payload (model declined to judge) | 0.065 | 0.016 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".