Time to get off the fence: The need for definitive international guidance on statistical analysis of ecotoxicity data
Bibliographic record
Abstract
The use of the no observed effect concentration (NOEC) and lowest observed effect concentration (LOEC) in ecotoxicology has been consistently criticized for over 30 years. A search of the literature from the past 30 years found 22 articles challenging the validity and/or appropriateness of NOEC/LOEC data compared to only one in defense of such data. Notwithstanding this compelling weight of evidence, the NOEC and LOEC remain commonly published measures of toxicity from ecotoxicological studies. In this article we argue that the major reason for the continued generation and publication of NOEC/LOEC data is that key government and intergovernmental organizations have been "sitting on the fence" on the issue for more than a decade. Although most key environmental quality guideline, toxicity testing, and associated guidance documents have now recognized the limitations of NOEC/LOEC data, to date no such document or standard toxicity test method has formally ceased recommending or providing guidance on the generation of such data. This is a problem because it is these very guidance documents and test methods that regulatory agencies demand be used by industry for regulatory activities, and on which commercial testing facilities attain and maintain their testing accreditation. Consequently, there will be little impetus for change to statistical analysis practices unless changes to the key guidance documents and test methods necessitate it. Although some progress on this has been made (e.g., in Canada, Australia and New Zealand), there needs to be stronger and universal action to ensure NOEC/LOEC data are no longer generated.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.006 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".