A statistical approach for the assessment and redesign of the Nile Delta drainage system water-quality-monitoring locations
Bibliographic record
Abstract
There are several deficiencies in the statistical approaches proposed in the literature for the assessment and redesign of surface water-quality-monitoring locations. These deficiencies vary from one approach to another, but generally include: (i) ignoring the attributes of the basin being monitored; (ii) handling multivariate water quality data sequentially rather than simultaneously; (iii) focusing mainly on locations to be discontinued; and (iv) ignoring the reconstitution of information at discontinued locations. In this paper, a methodology that overcomes these deficiencies is proposed. In the proposed methodology, the basin being monitored is divided into sub-basins, and a hybrid-cluster analysis is employed to identify groups of sub-basins with similar attributes. A stratified optimum sampling strategy is then employed to identify the optimum number of monitoring locations at each of the sub-basin groups. An aggregate information index is employed to identify the optimal combination of locations to be discontinued. The proposed approach is applied for the assessment and redesign of the Nile Delta drainage water quality monitoring locations in Egypt. Results indicate that the proposed methodology allows the identification of (i) the optimal combination of locations to be discontinued, (ii) the locations to be continuously measured and (iii) the sub-basins where monitoring locations should be added. To reconstitute information about the water quality variables at discontinued locations, regression, artificial neural network (ANN) and maintenance of variance extension (MOVE) techniques are employed. The MOVE record extension technique is shown to result in a better performance than regression or ANN for the estimation of information about water quality variables at discontinued locations.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.005 | 0.019 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.004 | 0.003 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".