Testing against “normal” with environmental data
Bibliographic record
Abstract
Abstract Normal ranges are some fraction of a reference distribution deemed to represent an expected condition, typically 95%. They are frequently used as the basis for generic criteria for monitoring programs designed to test whether a sample is outside of “normal,” as in reference-condition approach studies. Normal ranges are also the basis for criteria for more classic environmental effects monitoring programs designed to detect differences in mean responses between reference and exposure areas. Limits on normal ranges are estimated with error that varies depending largely on sample size. Direct comparison of a sample or a mean to estimated limits of a normal range will, with some frequency, lead to incorrect conclusions about whether a sample or a mean is inside or outside the normal range when the sample or the mean is near the limit. Those errors can have significant costs and risk implications. This article describes tests based on noncentral distributions that are appropriate for quantifying the likelihood that samples or means are outside a normal range. These noncentral tests reverse the burden of evidence (assuming that the sample or mean is at or outside normal), and thereby encourage proponents to collect more robust sample sizes that will demonstrate that the sample or mean is not at the limits or beyond the normal range. These noncentral equivalence and interval tests can be applied to uni- and multivariate responses, and to simple (e.g., upstream vs downstream) or more complex (e.g., before vs after, or upstream vs downstream) study designs. Statistical procedures for the various tests are illustrated with benthic invertebrate community data collected as part of the Regional Aquatics Monitoring Program (RAMP) in the vicinity of oil sands operations in northern Alberta, Canada. An Excel workbook with functions and calculations to carry out the various tests is provided in the online Supplemental Data. Integr Environ Assess Manag 2017;13:188–197. © 2016 SETAC. Key Points The article provides clarity on appropriate statistical procedures for testing whether a single observation or a mean falls outside a normal range of variation for a reference distribution. Three statistical procedures are described that test whether an observation or mean falls outside a normal range, and each uses noncentral distributions of test statistics. The procedures are illustrated with real data from simplified examples, in order to give readers a template on which to use the calculations in their work. An Excel workbook is provided with functions and calculations that can be used to compute the various tests.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".