Beyond Words: Enhancing Clinical Guideline Comprehension With Icons
Bibliographic record
Abstract
Abstract Background The Grading Recommendations, Assessment, Development, and Evaluations (GRADE) framework is widely applied in clinical guidelines to facilitate transparent evidence evaluation. While developing Infectious Diseases Society of America (IDSA) guidelines on the management of patients with coronavirus disease 2019 (COVID-19), panel members suggested developing and implementing a visual aid to enable quicker identification of key information by providers at bedside seeking guidance. Methods We conducted a mixed-methods study evaluating the usability of a newly designed infographic/icon using a survey and focus groups. The survey incorporated a simulated COVID-19 IDSA guideline with and without the icon, followed by comprehension questions. Focus group discussions provided qualitative feedback on the GRADE methodology and icon usability. Results The survey was returned by 289 health care providers. There was no statistical difference in the correct response rates between icon-aided and non-icon-aided guideline questions (McNemar's chi-square test, P > .1 for both questions). Interactions with the icon notably increased the time taken and number of clicks required to respond to the first question (Wilcoxon signed-rank test, P < .01). In contrast, response time did not differ between versions for the second question (P = .38). Most subjects (85%) indicated that the icon improved the readability of the guidelines. A focus group follow-up suggested alternative designs for the icon. Conclusions This study highlights the promise of iconography in clinical guidelines, although the specific icons tested did not measurably improve usability metrics. Future research should focus on icon design and testing within a formal usability framework, considering the impact of GRADE language on user experience.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; both teacher heads agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".