A code of practice for the conduct of systematic reviews in toxicology and environmental health research (COSTER)
Bibliographic record
Abstract
<strong>Background</strong>: There are several standards which make explicit a consensus view on sound practice in systematic reviews (SRs) for the medical sciences. Until now, no equivalent standard has been published for SRs which focus on human health risks posed by exposure to environmental challenges, chemical or otherwise. <strong>Objectives</strong>: To develop an expert, cross-sector consensus on a core set of requirements for sound practice in planning and conducting a SR in the environmental health sciences. <strong>Methods</strong>: A draft set of requirements was derived from two existing standards for SRs in biomedicine and discussed at an international workshop of 33 participants from government, industry, non-government organisations, and academia. The guidance was revised over six follow-up webinars and several rounds of email feedback, until there was group consensus that a comprehensive framework for the planning and conduct of high-quality environmental health SRs had been articulated. <strong>Results</strong>: The Conduct of Systematic Reviews in Toxicology and Environmental Health Research (COSTER) standard is a code of practice consisting of 70 requirements across eight performance domains, representing the consensus view of a diverse group of experts as to what constitutes “sound and good” practice in the conduct of environmental health SRs. <strong>Discussion</strong>: COSTER provides a set of sound-practice requirements which, if followed, should facilitate the production of credible, high-value SRs of environmental health evidence. COSTER clarifies sound and good practice in a number of controversial aspects of SR conduct, providing requirements relating to management of conflicts of interest, inclusion of grey literature, and protocol registration and publication. Not all of the practices are yet commonplace, but environmental health SRs would benefit from their introduction. Some aspects of SR, such as assessment of external validity at the level of individual study, are not yet sufficiently developed for consensus on sound practice to be achieved.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.022 | 0.012 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".