Assessing Space, Time, and Remediation Contribution to Soil Pollutant Variation near the Detection Limit Using Hurdle Models to Account for a Large Proportion of Nondetectable Results
Bibliographic record
Abstract
Many emerging, and some legacy, pollutants pose risks to humans and ecosystems near the detection limits (DL) of existing analytical systems. As a result, site assessments and management options are often presented with data sets that are sparse, highly skewed, and left-censored. Existing analysis methods are unable to differentiate effects of treatment from covariates, such as space, obscuring influences of site management. As a case study, we computed the mean and variance of censored soil benzene data across four sites over a three year period by gamma distribution with a maximum likelihood. Further, a combined hurdle model to accommodate left-censored concentrations was applied to analyze factors affecting benzene variation. This approach allowed us to assess the success and spatial dependency of a biostimulatory solution in reducing benzene concentrations at very low concentrations. Benzene concentrations decreased due to the addition of biostimulatory solution and spatial effects, but the detection of soil benzene after biostimulation was highly spatially dependent. By combining computed values for censored observations estimated by the hurdle-gamma model and uncensored observations, we can get the pseudocomplete data sets. The combined model is ideally suited to evaluate existing and emerging pollutants, that pose risks to humans and ecosystems but are typically at or near analytical detection limits.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".