Adaptive Surrogate Likelihood Function for Blended Hydrologic Models
Bibliographic record
Abstract
This abstract introduces a recipe for an adaptive general likelihood function and its application in the Bayesian epistemology of model parameters and structure uncertainty. The proposed methodology focuses on a special class of likelihood function, hereinafter mentioned as adaptive general likelihood function (AGL), which require a minimum priori assumptions/knowledge about the model residuals. The goal of the AGL is to characterize the model residuals independently from the inference framework in order to avoid incorrectly posterior estimation as a result of jointly inferencing of model and error model parameters. Mathematically, AGL is structured with a mixture of gaussian distributions joined with a first order autoregressive model, account for error model shape and autocorrelation respectively. To assess the AGL application, it is benchmarked with a formal likelihood function formulated by Schoups and Vrugt (2010) and evaluated for 24 Camels basins where the blended model has been deterministically applied with success (Chlumsky et al. 2022). Both approaches are compared with the residual’s empirical distributions using various statistical tests. The model used here is a blended hydrologic model introduced by Mai et al., (2021) which is a class of hydrologic models constructed by averaging (blending) various process options at the process flux level. This blending means calibration of the model functions to identify traditionally calibrated model process parameters as well as the weights utilized to average multiple process options. The model is deployed in the Raven hydrologic framework (Craig et al., 2020) and simultaneously both processes weights and parameters were calibrated deterministically for both high flows and low flows using PADDS algorithm (Asadzadeh and Tolson, 2013). This multi-objective calibration yields a suite of sample of calibrated blended models which is then utilized for error model development and testing. The tests results indicated a statistically comparable performance for both methods for t-distributed residuals highly skewed and long-tailed residual errors which are apparent in many hydrologic model residuals. Finally, to disjoin the epistemic Bayesian inference framework from the error model parameters, an epsilon-support vector regression (eps-SVR) is deterministically trained as a surrogate model to map the structural/parametric variability to residual error model parameters. The eps-SVR calibration performance metrics indicated high quality of surrogate for training set indicating promising performance.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".