Poorly-informative priors in geotechnical risk analysis
Bibliographic record
Abstract
Quantitative risk analysis has become a common part of geotechnical engineering. In domains such as dam safety, seismic hazard assessment, and flood damage reduction it has come to depend to a large extent on personalized probabilities in the Ramsey-deFinetti-Savage sense derived by quantifying engineering judgment. From a Bayesian view, quantified judgments principally manifest in prior probabilities, which may be informed by prior information, but which also may be poorly- or un-informed and based on subjective experience. The choice of Likelihood function within the Bayesian context also introduces personalistic uncertainty, but that is infrequently considered. We use the term poorly-informative to differentiate from the non-informative prior in the Jeffreys sense. As the field becomes more receptive to risk-informed thinking, the question of how to quantify and calibrate judgment has become more pressing. How do we quantify priors in a way that is aligned with reality? How much difference does vagueness in the prior make in engineering predictions? Do we weight different experts’ probabilities differently? We now have four decades of experience in attempting to quantify geotechnical judgment in the aleatory domain where chance is dominant and in the epistemic domain where inadequate knowledge is dominant. This experience is reflected upon to draw lessons and to create workable suggestions for practice. The paper principally draws on experience with risk analysis in dam safety. How well-calibrated is an expert when assigning probabilities to parameters or to events in the world? Since probabilities in the Bayesian sense are degrees of belief, the assignment of probability is always correct to the extent that it accurately reflects an expert’s belief. Two people can assign different probabilities and both be “right.” Yet, if a consultant is hired for the purpose of contributing information from which to make decisions, one would like to know whether that expert’s beliefs are consistent with frequencies in the world. Is he or she calibrated? How can quantified expert opinion be validated considering ex post observations of engineering performance, especially failures? A quantitative Bayesian validation procedure is proposed based on the concept of expert-as-information in the sense of Morris and used to assess the credibility of experts. This is applied to how a decision-maker should ascribe credibility to an expert’s judgments when attempting to predict the performance of engineering designs.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.045 | 0.067 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.001 |
| Bibliometrics | 0.016 | 0.065 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.002 | 0.002 |
| Open science | 0.007 | 0.004 |
| Research integrity | 0.001 | 0.003 |
| Insufficient payload (model declined to judge) | 0.001 | 0.005 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; both teacher heads agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".