Recommendations for the conduct of systematic reviews in toxicology and environmental health research (COSTER)
Bibliographic record
Abstract
BACKGROUND: There are several standards that offer explicit guidance on good practice in systematic reviews (SRs) for the medical sciences; however, no similarly comprehensive set of recommendations has been published for SRs that focus on human health risks posed by exposure to environmental challenges, chemical or otherwise. OBJECTIVES: To develop an expert, cross-sector consensus view on a key set of recommended practices for the planning and conduct of SRs in the environmental health sciences. METHODS: A draft set of recommendations was derived from two existing standards for SRs in biomedicine and developed in a consensus process, which engaged international participation from government, industry, non-government organisations, and academia. The consensus process consisted of a workshop, follow-up webinars, email discussion and bilateral phone calls. RESULTS: The Conduct of Systematic Reviews in Toxicology and Environmental Health Research (COSTER) recommendations cover 70 SR practices across eight performance domains. Detailed explanations for specific recommendations are made for those identified by the authors as either being novel to SR in general, specific to the environmental health SR context, or potentially controversial to environmental health SR stakeholders. DISCUSSION: COSTER provides a set of recommendations that should facilitate the production of credible, high-value SRs of environmental health evidence, and advance discussion of a number of controversial aspects of conduct of EH SRs. Key recommendations include the management of conflicts of interest, handling of grey literature, and protocol registration and publication. A process for advancing from COSTER's recommendations to developing a formal standard for EH SRs is also indicated.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.070 | 0.012 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.007 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; both teacher heads agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".