Interlaboratory evaluation of<i>Hyalella azteca</i>and<i>Chironomus tentans</i>short-term and long-term sediment toxicity tests
Bibliographic record
Abstract
Methods for assessing the long-term toxicity of sediments to Hyalella azteca and Chironomus tentans can significantly enhance the capacity to assess sublethal effects of contaminated sediments through multiple endpoints. Sublethal tests allow us to begin to understand the relationship between short-term and long-term effects for toxic sediments. We present an interlaboratory evaluation with long-term and 10-d tests using control and contaminated sediments in which we assess whether proposed and existing performance criteria (test acceptability criteria [TAC]) could be achieved. Laboratories became familiar with newly developed, long-term protocols by testing two control sediments in phase 1. In phase 2, the 10-d and long-term tests were examined with several sediments. Laboratories met the TACs, but results varied depending on the test organism, test duration, and endpoints. For the long-term tests in phase 1, 66 to 100% of the laboratories consistently met the TACs for survival, growth, or reproduction using H. azrteca, and 70 to 100% of the laboratories met the TACs for survival and growth, emergence, reproduction, and hatchability using C. tentans. In phase 2, fewer laboratories participated in long-term tests: 71 to 88% of the laboratories met the TAC for H. azteca, whereas 50 to 67% met the TAC for C. tentans. In the 10-d tests with H. azteca and C. tentans, 82 and 88% of the laboratories met the TAC for survival, respectively, and 80% met the TAC for C. tentans growth. For the 10-d and long-term tests, laboratories predicted similar toxicity. Overall, the interlaboratory evaluation showed good precision of the methods, appropriate endpoints were incorporated into the test protocols, and tests effectively predicted the toxicity of sediments.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.023 | 0.019 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.002 | 0.001 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.002 | 0.003 |
| Research integrity | 0.002 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".