Predicted no effect concentration derivation as a significant source of variability in environmental hazard assessments of chemicals in aquatic systems: An international analysis
Bibliographic record
Abstract
Environmental hazard assessments for chemicals are carried out to define an environmentally "safe" level at which, theoretically, the chemical will not negatively affect any exposed biota. Despite this common goal, the methodologies in use are very diverse across different countries and jurisdictions. This becomes particularly obvious when international scientists work together on documents with global scope, e.g., in the World Health Organization (WHO) International Program on Chemical Safety. In this article, we present a study that describes the extent of such variability and analyze the reasons that lead to different outcomes in deriving a "safe level" (termed the predicted no effect concentration [PNEC] throughout this article). For this purpose, we chose 5 chemicals to represent well-known substances for which sufficient high-quality aquatic effects data were available: ethylene glycol, trichloroethylene, nonylphenol, hexachlorobenzene, and copper (Cu). From these data, 2 data sets for each chemical were compiled: the full data set, that contained all information from selected peer-review sources, and the base data set, a subsample of the full set simulating limited data. Scientists from the European Union (EU), United States, Canada, Japan, and Australia independently carried out hazard assessments for each of these chemicals using the same data sets. Their reasoning for key study selection, use of assessment factors, or use of probabilistic methods was comprehensively documented. The observed variation in the PNECs for all chemicals was up to 3 orders of magnitude, and this was not simply due to obvious factors such as the size of the data set or the methodology used. Rather, this was due to individual decisions of the assessors within the scope of the methodology used, especially key study selection, acute versus chronic definitions, and size of assessment factors. Awareness of these factors, together with transparency of the decision-making process, would be necessary to minimize confusion and uncertainty related to different hazard assessment outcomes, particularly in international documents. The development of a "guideline on transparency in decision-making" ensuring the decision-making process is science-based, understandable, and transparent, may therefore be a promising way forward.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.004 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".