Comparison of Artificial Neural network (ANN) models and conventional GIS methods to interpolate high-resolution maps of sediment contamination in the Detroit River
Bibliographic record
Abstract
Deep learning Artificial Neural Network (DNN) models that included flow velocity, flow directionality, various positional inputs, and sediment total organic carbon (TOC) content as predictor variables were compared to conventional GIS interpolation for predicting sediment contamination of zinc (Zn), mercury (Hg), polychlorinated biphenyls (PCBs), and polycyclic aromatic hydrocarbons (PAHs) in the Detroit River Area of Concern (AOC). A sediment chemistry dataset of 500–850 sampling stations was split into training and validation data across 1000 splits. Universal kriging provided the best interpolation fits relative to other conventional GIS methods including inverse distance weighting and spline fitting. Parsimony-optimized ANN models explained more variation (43.0–58.4%) in validation data splits compared to kriging (19.7– 40.2%) and other GIS approaches. Kriging tended to overestimate contaminant concentrations, especially for areas containing low Hg and PCB levels. ANN performance was improved by adjusting predictions for U.S. restoration areas, non-U.S. restoration areas, and Canadian zones. High-resolution contaminant maps were generated for 9668 cells throughout the AOC. The two approaches showed differences in predicted sediment contamination distribution, with kriging predicting lower river-wide Zn and Hg contamination and higher PCB and PAH contamination. ANNs generated more discrete contaminant hot zones, indicating strong nearshore gradients along the U.S. shoreline and low concentrations in fast moving navigation channels. Given the superior performance of ANNs, this approach was considered a superior interpolation method, with potential applications for guiding future sediment restoration initiatives in the AOC.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.003 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".