A Model Selection Procedure for Stream Re-Aeration Coefficient Modelling
Bibliographic record
Abstract
Model selection is finding wide applications in a lot of modelling and environmental problems. However, applications of model selection to re-aeration coefficient studies are still limited. The current study explores the use of model selection in re-aeration coefficient studies by combining several suggestions from numerous authors on the interpretation of data regarding re-aeration coefficient modelling. The model selection procedure applied in this research made use of Akaike information criteria, measures of agreement such as percent bias (PBIAS), Nash-Sutcliffe Efficiency (NSE) and RMSE observation Standard deviation Ratio (RSR) and gragh analysis in selecting the best performing model. An algorithm prescribing a generic model selection procedure was also provided. Out of ten candidates models used in this study, the O’Connor and Dobbins (1958) model emerged as the top performing model in its application to data collected from River Atuwara in Nigeria. The suggested process could save software and model developers lots of time and resources, which would otherwise be spent in investigating and developing new models. The procedure is also ideal in selecting a model in situations where there is no overwhelming support for any particular model by observed data.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.007 | 0.015 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.003 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.002 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.004 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".