Application of optimization techniques to water quality monitoring designs
Bibliographic record
Abstract
The issues of possible improvements, increased efficiency and/or optimization of monitoring systems in general, and monitoring designs, in particular, attract the attention of researchers for years.Traditionally, these issues are addressed using expert knowledge and heuristic approaches to monitoring system development under the budgetary constraints.Application of formal techniques for these purposes looks appealing since it may validate suggested procedures or justify expenses required for data collection.The paper describes an approach to the development of sampling programs as a solution of the operation research model articulated in terms of the cost-effectiveness analysis.The effectiveness of a sampling program is described through the uncertainty of the estimates obtained from water quality data collected in accordance with the sampling program.Since several water quality indicators, with different temporal and/or spatial variability, are determined from the same water sample, monitoring designs for these indicators must be compromised in a way to ensure a required level of efficiency for all indicators being detected.The proposed approach is based on the operation research model which minimizes the total number of water samples been collected over an investigated period of time under the condition that the uncertainty of an estimate derived from the monitoring data is kept below an acceptable level.The efficient monitoring designs are determined as the solutions of this model.Since concentrations of water constituents exhibit different variability, the numbers of observations required to achieve the same uncertainty level in their estimates vary significantly, even for those water constituents whose concentrations are derived from the same grab water sample.In order to make a practically meaningful recommendation on the frequencies of observations, it is necessary to comprise temporal monitoring designs for all water quality indicators from the same water sample.Given that concentrations of these parameters form under common hydrological and climatic conditions, it is reasonable to assume that series of concentrations are somehow related.It had been shown that if such dependencies are detected, they can be used to significantly reduce the total number of observations required for water quality assessment.The proposed approach has been tested on observation data collected on the Humber River (Ontario, Canada).The major ions, namely, calcium, carbon, magnesium, and potassium have been selected for the study.Since monitoring data can be used for various purposes, simple random designs supporting the evaluation of basic statistics of the investigated water quality indicators are preferable.The relationships between concentrations of the investigated water constituents were described by linear regression models.These models were used in the proposed operation research model to obtain efficient monitoring designs supporting estimation of water quality indicators with a given level of uncertainty.These designs are common for all investigated water quality indicators detected from the same water sample at the Old Mill Road cross-section of the Humber River.The proposed operation research model can be applied to tiered monitoring systems where water quality indicators of interest are split in two sets: core and supplemental, according to their importance for a given site with different accuracy requirements.The designs common for all water quality indicators measured at the given site may result in a higher number of observations.Depending on the desired level of accuracy, it may lead to daily sample collection.The proposed approach may help to develop efficient monitoring designs with the reasonable cost of sampling by considering subsets of the water quality indicators.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.006 | 0.016 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".