Hydrogeophysical data integration through Bayesian Sequential Simulation with log-linear pooling
Bibliographic record
Abstract
SUMMARY Bayesian sequential simulation (BSS) is a geostastistical technique, which uses a secondary variable to guide the stochastic simulation of a primary variable. As such, BSS has proven significant promise for the integration of disparate hydrogeophysical data sets characterized by vastly differing spatial coverage and resolution of the primary and secondary variables. An inherent limitation of BSS is its tendency to underestimate the variance of the simulated fields due to the smooth nature of the secondary variable. Indeed, in its classical form, the method is unable to account for this smoothness because it assumes independence of the secondary variable with regard to neighbouring values of the primary variable. To overcome this limitation, we have modified the Bayesian updating with a log-linear pooling approach, which allows us to account for the inherent interdependence between the primary and the secondary variables by adding exponential weights to the corresponding probabilities. The proposed method is tested on a pertinent synthetic hydrogeophysical data set consisting of surface-based electrical resistivity tomography (ERT) data and local borehole measurements of the hydraulic conductivity. Our results show that, compared to classical BSS, the proposed log-linear pooling method using equal constant weights for the primary and secondary variables enhances the reproduction of the spatial statistics of the stochastic realizations, while maintaining a faithful correspondence with the geophysical data. Significant additional improvements can be achieved by optimizing the choice of these constant weights. We also explore a dynamic adaptation of the weights during the course of the simulation process, which provides valuable insights into the optimal parametrization of the proposed log-linear pooling approach. The results corroborate the strategy of selectively emphasizing the probabilities of the secondary and primary variables at the very beginning and for the remainder of the simulation process, respectively.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".