It is time for changes in the analysis of whole effluent toxicity data
Bibliographic record
Abstract
The whole effluent toxicity (WET) program in the United States, Canada, and other countries typically requires multi concentration testing of effluents. While multiconcentration testing of chemicals is desirable for regulatory and scientific reasons, we believe this requirement is not as efficient for evaluating effluent compliance in a WET program. The key regulatory question of concern is whether an effluent is toxic or not, which is best answered statistically using a hypothesis approach, not a point estimate approach. However, the traditional hypothesis approach currently recommended does not reward high within-test precision. This report describes the need for 3 specific changes in the analysis of WET compliance data that we believe would yield a more robust WET regulatory program: (1) restate the null hypothesis so that test power is associated with demonstrating that the effluent is not toxic, (2) use USEPA's Test of Significant Toxicity (based on the noninferiority approach) to identify unacceptable toxicity as well as acceptable effects with a high probability, and (3) evaluate only the test control and the critical concentration of concern (e.g., instream waste concentration). We demonstrate that instituting these 3 changes would provide: Positive incentives for permittees to produce high-quality WET data, a transparent analysis approach in which the permittee could have greater control over regulatory decisions based on test results, and potentially a less expensive testing program because fewer effluent concentrations need to be examined within a test. As a result, WET test frequency could be increased for the same cost as current testing programs while providing greater representativeness of effluent quality.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".