MétaCan
Menu
Back to cohort
Record W2917647106 · doi:10.1002/qre.1040

Discussion (3): Jones–Johnson Paper

2009· article· en· W2917647106 on OpenAlexaff
Jason L. Loeppky, Brian J. Williams

Bibliographic record

VenueQuality and Reliability Engineering International · 2009
Typearticle
Languageen
FieldComputer Science
TopicAdvanced Multi-Objective Optimization Algorithms
Canadian institutionsOkanagan University CollegeUniversity of British Columbia, Okanagan CampusUniversity of British Columbia
Fundersnot available
KeywordsComputer science

Abstract

fetched live from OpenAlex

We commend Jones and Johnson for providing a clear and concise introduction to the quickly expanding field of statistical design and analysis of computer experiments.Our own experiences confirm that computer models are widely used in many areas of science and engineering, and the need to understand the performance of these models is critical.This is particularly true in light of the fact that with improving model fidelity, they are increasingly used to certify complex engineering systems with decreasing reliance on expensive physical experiments.As the authors indicate, many computer models are sufficiently complex that only a small budget of model runs is allowed for any given application.Therefore, concepts of statistical experiment design become relevant for the purpose of intelligently selecting runs to inform the development of a statistical surrogate for model output-often referred to as an emulator-that will serve as the basis for statistical inference.There are several practical issues with emulating any computer model.Is the standard Gaussian Process (GP) model a good choice for general applications?What mean and covariance structure should one choose?How should the budget of runs be expended?The authors addressed these issues effectively in their article.We provide some additional perspective in what follows.Extensive literature (see Sacks et al. 1 , Santner et al. 2 ) and experience suggest that the GP model is an ideal candidate for building an emulator.Ben-Ari and Steinberg 3 conducted an extensive simulation study comparing the GP model with a large class of competing models and found that GP-based emulation performs well in many situations.There are several practical considerations that must be addressed when using the GP model.Arguably the most important is how many runs are needed to adequately emulate the computer model.Loeppky et al. 4 argue that the often quoted rule of 'n = 10d' (i.e. 10 model runs per dimension) generally provides sufficient information for emulation.The design chosen for the example of this article comes close to attaining this target, at 8.75d.In addition to run size considerations, it is important to calculate diagnostics that directly assess the quality of model fit.The root mean square error (RMSE) is an obvious criterion; however, holdout samples are often unavailable in practice.In such cases the cross-validated (CV)-RMSE (see Welch et al. 5 ) and the individual CV residuals are useful.The remainder of this discussion is structured to draw attention to additional connections between traditional response surface methodology (RSM) and analysis of computer experiments using the GP model.In particular, we focus on two basic components of response surface methods (see Box and Wilson 6 ) for which recent developments have made analogues available for analysis of computer experiments: sensitivity analysis and sequential optimization.Sensitivity analysis refers to measuring the impact of input variations on output uncertainty.In particular, output uncertainty can be decomposed into main and interaction effects analogous to traditional analysis of variance (ANOVA), and sensitivity indices measuring the contribution of these individual effects to the total output variance can be computed (see Saltelli et al. 7 , Oakley and O'Hagan 8 , Schonlau and Welch 9 ).Sequential optimization of computer models based on the expected improvement criteria has proven efficient and effective (see Jones et al. 10 ).These optimization algorithms are global in the sense that they explore regions of the input space in which prediction is poor (potential for optima), while focusing in on regions of space containing optima with high probability.Goodness-of-fit diagnostics, sensitivity analysis and sequential optimization are explored with two analyses of the example using the F-quantile function presented in this article.The first analysis (referred to as

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.008
metaresearch head score (Gemma)0.030
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: Not applicable
GenreCandidate signal: Commentary · Consensus signal: Commentary
Teacher disagreement score0.102
Threshold uncertainty score0.341

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0080.030
Meta-epidemiology (narrow)0.0010.000
Meta-epidemiology (broad)0.0010.001
Bibliometrics0.0010.001
Science and technology studies0.0040.002
Scholarly communication0.0060.005
Open science0.0020.003
Research integrity0.0150.008
Insufficient payload (model declined to judge)0.1020.050

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.011
GPT teacher head0.283
Teacher spread0.272 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designNot applicable
Domainnot available
GenreCommentary

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2009
Admission routes1
Has abstractyes

Explore more

Same venueQuality and Reliability Engineering InternationalSame topicAdvanced Multi-Objective Optimization AlgorithmsFrench-language works237,207