Tips for Teachers of Evidence-based Medicine: Making Sense of Diagnostic Test Results Using Likelihood Ratios
Bibliographic record
Abstract
Now is an exciting time to be or become a diagnostician. More diagnostic tests, including portions of the medical interview and physical examination, are being studied rigorously for their accuracy, precision, and usefulness in practice,1,2 and this research is increasingly being systematically reviewed and synthesized.3,4 Diagnosticians are gaining increasing access to this research evidence, raising hope that this knowledge will inform their diagnostic decisions and improve their patients’ clinical outcomes.5 For patients to benefit fully from this accumulating knowledge, the diagnosticians serving them must be able to reason probabilistically, to understand how test results can revise disease probability to confirm or exclude disorders, and to integrate this reasoning with other types of knowledge and diagnostic thinking.6–8 Yet, clinicians encounter several barriers when trying to integrate research evidence into clinical diagnosis.9 Some barriers involve difficulties in understanding and using the quantitative measures of tests’ accuracy and discriminatory power, including sensitivity, specificity, and likelihood ratios (LRs).9,10 We have noticed that LRs are particularly troubling to many learners at first, and we have wondered if this is because of the way they have been taught. Stumbling blocks can arise in several places when learning LRs: the names and formulae themselves can be intimidating; the arithmetic functions can be mystifying when attempted all at once; if two levels of test results are taught first, learners can have difficulty ‘stretching’ to multiple levels; and if disease probability is framed in odds terms (to directly multiply the odds by the likelihood ratio), learners can misunderstand why and how this conversion is done. Other stumbling blocks may occur as well. Other authors have described various approaches to helping clinicians understand LRs.11–16 In this article, we describe two additional approaches to help clinical learners understand how LRs describe the discriminatory power of test results. Whereas we mention other concepts such as pretest and posttest probability, full treatment of those subjects is beyond the scope of this article. These approaches were developed by experienced teachers of evidence-based medicine (EBM) and were refined over years of teaching practice. These tips have also been field-tested to double-check the clarity and practicality of these descriptions, as explained in the introductory article of this series.17 To help the reader envision these teaching approaches, we present sequenced advice for teachers in plain text, coupled with sample words to speak, in italics. These scripts are meant to be interactive, which means that teachers should periodically check in with the learners for their understanding and that teachers should try other ways to explain the ideas if the words we have suggested do not “click.” We present them in order from shorter to longer; however, because these 2 scripts cover the same general content, we encourage teachers to use either or both in an order that best fits their setting and learners.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.069 | 0.458 |
| Meta-epidemiology (narrow) | 0.003 | 0.003 |
| Meta-epidemiology (broad) | 0.004 | 0.003 |
| Bibliometrics | 0.009 | 0.004 |
| Science and technology studies | 0.002 | 0.011 |
| Scholarly communication | 0.010 | 0.034 |
| Open science | 0.007 | 0.007 |
| Research integrity | 0.018 | 0.056 |
| Insufficient payload (model declined to judge) | 0.025 | 0.021 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".