Modified Delphi survey for the evidence summarisation of patient decision aids: Study protocol
Bibliographic record
Abstract
INTRODUCTION: Information included in a patient decision aid (PDA) can significantly influence patients' decisions and is, therefore, expected to be evidence-based and rigorously selected and summarised. PDA developers have not yet agreed on a standardised process for the selection and summarisation of the supporting evidence. We intend to generate consensus on a process (and related steps and criteria) for selecting and summarising evidence for PDAs using a modified Delphi survey. METHODS AND ANALYSIS: We will develop an evidence summarisation process specific to PDA development by using a consensus-based Delphi approach, surveying international experts and stakeholders with two to three rounds. To increase generalisability and acceptability, we will distribute the survey to the following stakeholder groups: PDA developers, researchers with expertise in shared decision making, PDA development and evidence summarisation, members of the International Patient Decision Aids Standards (IPDAS) collaboration, policy makers with expertise in PDA certification and patient stakeholder groups. For each criterion, if at least 80% of survey participants rank the criterion as most important/least important, we will consider that consensus has been achieved. ETHICS AND DISSEMINATION: It is critical for PDAs to have accurate and trustworthy evidence-based information about the risks and benefits of health treatments and tests, as these decision aids help patients make important choices. We want to generate consensus on an approach for selecting and summarising the evidence included in PDAs, which can be widely implemented by PDA developers. Dartmouth College's Committee for the Protection of Human Subjects approved this protocol. We will publish our results in a peer reviewed journal.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.020 | 0.007 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.002 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".