MétaCan
Menu
Retour à la cohorte
Enregistrement W2917647106 · doi:10.1002/qre.1040

Discussion (3): Jones–Johnson Paper

2009· article· en· W2917647106 sur OpenAlexaff
Jason L. Loeppky, Brian J. Williams

Notice bibliographique

RevueQuality and Reliability Engineering International · 2009
Typearticle
Langueen
DomaineComputer Science
ThématiqueAdvanced Multi-Objective Optimization Algorithms
Établissements canadiensOkanagan University CollegeUniversity of British Columbia, Okanagan CampusUniversity of British Columbia
Organismes subventionnairesnon disponible
Mots-clésComputer science

Résumé

récupéré en direct d'OpenAlex

We commend Jones and Johnson for providing a clear and concise introduction to the quickly expanding field of statistical design and analysis of computer experiments.Our own experiences confirm that computer models are widely used in many areas of science and engineering, and the need to understand the performance of these models is critical.This is particularly true in light of the fact that with improving model fidelity, they are increasingly used to certify complex engineering systems with decreasing reliance on expensive physical experiments.As the authors indicate, many computer models are sufficiently complex that only a small budget of model runs is allowed for any given application.Therefore, concepts of statistical experiment design become relevant for the purpose of intelligently selecting runs to inform the development of a statistical surrogate for model output-often referred to as an emulator-that will serve as the basis for statistical inference.There are several practical issues with emulating any computer model.Is the standard Gaussian Process (GP) model a good choice for general applications?What mean and covariance structure should one choose?How should the budget of runs be expended?The authors addressed these issues effectively in their article.We provide some additional perspective in what follows.Extensive literature (see) and experience suggest that the GP model is an ideal candidate for building an emulator.Ben-Ari and Steinberg 3 conducted an extensive simulation study comparing the GP model with a large class of competing models and found that GP-based emulation performs well in many situations.There are several practical considerations that must be addressed when using the GP model.Arguably the most important is how many runs are needed to adequately emulate the computer model.Loeppky et al. 4 argue that the often quoted rule of 'n = 10d' (i.e. 10 model runs per dimension) generally provides sufficient information for emulation.The design chosen for the example of this article comes close to attaining this target, at 8.75d.In addition to run size considerations, it is important to calculate diagnostics that directly assess the quality of model fit.The root mean square error (RMSE) is an obvious criterion; however, holdout samples are often unavailable in practice.In such cases the cross-validated (CV)-RMSE (see Welch et al. 5 ) and the individual CV residuals are useful.The remainder of this discussion is structured to draw attention to additional connections between traditional response surface methodology (RSM) and analysis of computer experiments using the GP model.In particular, we focus on two basic components of response surface methods (see Box and Wilson 6 ) for which recent developments have made analogues available for analysis of computer experiments: sensitivity analysis and sequential optimization.Sensitivity analysis refers to measuring the impact of input variations on output uncertainty.In particular, output uncertainty can be decomposed into main and interaction effects analogous to traditional analysis of variance (ANOVA), and sensitivity indices measuring the contribution of these individual effects to the total output variance can be computed (see Saltelli et al. 7 , Oakley and O'Hagan 8 , Schonlau and Welch 9 ).Sequential optimization of computer models based on the expected improvement criteria has proven efficient and effective (see Jones et al. 10 ).These optimization algorithms are global in the sense that they explore regions of the input space in which prediction is poor (potential for optima), while focusing in on regions of space containing optima with high probability.Goodness-of-fit diagnostics, sensitivity analysis and sequential optimization are explored with two analyses of the example using the F-quantile function presented in this article.The first analysis (referred to as

Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.

Comment cette classification a été obtenuedéplier

Prédiction machine sur la base complète

Imitation des enseignants

Ni prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.

score de la tête « metaresearch » (Codex)0,008
score de la tête « metaresearch » (Gemma)0,030
Version: metacan-v3-hybrid-931329e0061cStatut de validation: machine_predicted_unvalidated
Catégories candidatesaucune
Catégories consensuellesaucune
DomaineSignal candidat: aucune · Signal consensuel: aucune
Devis d'étudeSignal candidat: Sans objet · Signal consensuel: Sans objet
GenreSignal candidat: Commentaire · Signal consensuel: Commentaire
Score de désaccord entre enseignants0,102
Score d'incertitude au seuil0,341

Scores du classifieur distillé par catégorie (deux têtes)

CatégorieCodexGemma
Métarecherche0,0080,030
Méta-épidémiologie (sens strict)0,0010,000
Méta-épidémiologie (sens large)0,0010,001
Bibliométrie0,0010,001
Études des sciences et des technologies0,0040,002
Communication savante0,0060,005
Science ouverte0,0020,003
Intégrité de la recherche0,0150,008
Charge utile insuffisante (le modèle a refusé de juger)0,1020,050

Scores machine (provisoires)

Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.

Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.

Tête enseignante Opus0,011
Tête enseignante GPT0,283
Écart entre enseignants0,272 · la distance entre les deux têtes enseignantes sur ce seul travail
Statut de validationscore_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découle

Classification

machine, non validée

Prédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.

Les modèles n’ont appliqué aucune catégorie : rien dans la taxonomie ne correspondait à ce travail.
Devis d'étudeSans objet
Domainenon disponible
GenreCommentaire

Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».

En bref

Citations0
Publié2009
Routes d'admission1
Résumé présentoui

Explorer davantage

Même revueQuality and Reliability Engineering InternationalMême sujetAdvanced Multi-Objective Optimization AlgorithmsTravaux en français237 207