Communicating Uncertainty From Limitations in Quality of Evidence to the Public in Written Health Information: Protocol for a Web-Based Randomized Controlled Trial
Bibliographic record
Abstract
BACKGROUND: Uncertainty is integral to evidence-informed decision making and is of particular importance for preference-sensitive decisions. Communicating uncertainty to patients and the public has long been identified as a goal in the informed and shared decision-making movement. Despite this, there is little quantitative research on how uncertainty in health information is perceived by readers. OBJECTIVE: The objective of this study is to design an experiment to examine how different degrees of uncertainty (Q1) and different types of uncertainty (Q2) impact patients' perception of treatment effectiveness, the body of evidence, text quality, and hypothetical treatment intention. The experiment also examines whether there is an additive effect when multiple sources of uncertainty are communicated (Q3). METHODS: We developed 8 variations of a research summary set in a hypothetical scenario for a treatment decision in the context of tinnitus. These were modified only in the degree of uncertainty relating to the evidence of the presented treatment. We recruited members of the German public from a Web-based research panel and randomized them to one of 8 variations of the research summary to examine the 3 research questions. The trial was only open to the members of the research panel. The outcomes are perception of the effectiveness of the treatment (primary), certainty in the judgement of treatment effectiveness, perception of the body of evidence relating to the treatment, text quality, and decisional intention (secondary). Outcomes were self-assessed. We aimed to recruit 1500 participants to the trial. The recruitment and data collection was fully automated. Ethical approval was waivered by an ethics committee because of the negligible risk to participants. RESULTS: This protocol is retrospectively published in its original format. In the meantime, the trial was set up and the data collection was completed. Data collection was conducted in May 2018. A total of 1727 eligible panel members were enrolled. CONCLUSIONS: We aim to publish the results in a peer-reviewed journal by the end of 2019. In addition, results will be presented at conferences and disseminated among developers of guidance for the development of evidence-based health information and decision aids. TRIAL REGISTRATION: German Clinical Trials Register DRKS00015911; https://www.drks.de/drks_web/navigate.do? navigationId=trial.HTML&TRIAL_ID=DRKS00015911 (Archived by WebCite at http://www.webcitation.org/77zyZTGzk). INTERNATIONAL REGISTERED REPORT IDENTIFIER (IRRID): DERR1-10.2196/13425.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.101 | 0.118 |
| Meta-epidemiology (narrow) | 0.009 | 0.005 |
| Meta-epidemiology (broad) | 0.011 | 0.009 |
| Bibliometrics | 0.006 | 0.006 |
| Science and technology studies | 0.005 | 0.007 |
| Scholarly communication | 0.007 | 0.006 |
| Open science | 0.004 | 0.003 |
| Research integrity | 0.011 | 0.013 |
| Insufficient payload (model declined to judge) | 0.089 | 0.019 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".