MétaCan
Menu
Back to cohort
Record W4220820710 · doi:10.1136/bmjopen-2021-053246

Defining decision thresholds for judgments on health benefits and harms using the Grading of Recommendations Assessment, Development and Evaluation (GRADE) Evidence to Decision (EtD) frameworks: a protocol for a randomised methodological study (GRADE-THRESHOLD)

2022· article· en· W4220820710 on OpenAlexaff
Gian Paolo Morgano, Lawrence Mbuagbaw, Nancy Santesso, Feng Xie, Jan Brożek, Uwe Siebert, Antonio Bognanni, Wojtek Wiercioch, Thomas Piggott, Andrea Darzi, Elie A. Akl, Ilse M. Verstijnen, Elena Parmelli, Zuleika Saz‐Parkinson, Pablo Alonso‐Coello, Holger J. Schünemann

Bibliographic record

VenueBMJ Open · 2022
Typearticle
Languageen
FieldEconomics, Econometrics and Finance
TopicHealth Systems, Economic Evaluations, Quality of Life
Canadian institutionsMcMaster UniversityImpactCochrane
Fundersnot available
KeywordsMedicineGrading (engineering)GuidelineProtocol (science)Psychological interventionApplied psychologyConsistency (knowledge bases)Transparency (behavior)Research ethicsExternal validityEvidence-based medicineDecision aidsMedical educationAlternative medicinePsychologyNursingSocial psychologyComputer sciencePathologyArtificial intelligence

Abstract

fetched live from OpenAlex

INTRODUCTION: The Grading of Recommendations Assessment, Development and Evaluation (GRADE) and similar Evidence to Decision (EtD) frameworks require its users to judge how substantial the effects of interventions are on desirable and undesirable people-important health outcomes. However, decision thresholds (DTs) that could help understand the magnitude of intervention effects and serve as reference for interpretation of findings are not yet available.The objective of this study is an approach to derive and use DTs for EtD judgments about the magnitude of health benefits and harms. We hypothesise that approximate DTs could have the ability to discriminate between the existing four categories of EtD judgments (Trivial, Small, Moderate, Large), support panels of decision-makers in their work, and promote consistency and transparency in judgments. METHODS AND ANALYSIS: We will conduct a methodological randomised controlled trial to collect the data that allow deriving the DTs. We will invite clinicians, epidemiologists, decision scientists, health research methodologists, experts in Health Technology Assessment (HTA), members of guideline development groups and the public to participate in the trial. Then, we will investigate the validity of our DTs by measuring the agreement between judgments that were made in the past by guideline panels and the judgments that our DTs approach would suggest if applied on the same guideline data. ETHICS AND DISSEMINATION: The Hamilton Integrated Research Ethics Board reviewed this study as a quality improvement study and determined that it requires no further consent. Survey participants will be required to read a consent statement in order to participate in this study at the beginning of the trial. This statement reads: You are being invited to participate in a research project which aims to identify indicative DTs that could assist users of the GRADE EtD frameworks in making judgments. Your input will be used in determining these indicative thresholds. By completing this survey, you provide consent that the anonymised data collected will be used for the research study and to be summarised in aggregate in publication and electronic tools. PROTOCOL REGISTRATION NUMBER: NCT05237635.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.122
metaresearch head score (Gemma)0.011
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch, Science and technology studies
Consensus categoriesMetaresearch
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Simulation or modeling · Consensus signal: none
GenreCandidate signal: Protocol · Consensus signal: Protocol
Teacher disagreement score0.407
Threshold uncertainty score1.000

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.1220.011
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0010.000
Bibliometrics0.0000.000
Science and technology studies0.0020.000
Scholarly communication0.0000.000
Open science0.0010.001
Research integrity0.0000.000
Insufficient payload (model declined to judge)0.0000.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.868
GPT teacher head0.666
Teacher spread0.202 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; both teacher heads agree on what is shown here.

Study designSimulation or modeling
Domainnot available
GenreProtocol

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations46
Published2022
Admission routes1
Has abstractyes

Explore more

Same venueBMJ OpenSame topicHealth Systems, Economic Evaluations, Quality of LifeFrench-language works237,207