Arming the Bayesian Physician to Rule Out Pulmonary Embolism: Using Evidence‐based Diagnostics to Combat Overtesting
Bibliographic record
Abstract
Overtesting and its downstream consequences (overdiagnosis and overtreatment) have recently become targets for both patients and policy-makers.1, 2 In an era of limited financial resources for medical costs at a societal level, proactive responses to overtesting include the national Choosing Wisely campaign and other similar local and regional efforts to define overused tests upon which we should be focusing a spotlight.3, 4 Within emergency medicine (EM), the potentially overused tests used to evaluate patients with suspected pulmonary embolism (PE) are prime targets for groups like “Preventing Overdiagnosis,”5 a collaborative effort between Consumer Reports, The Dartmouth Institute, the British Medical Journal, and Monash University. While PE is a potentially life-threatening disorder, it often has an atypical presentation that mimics multiple other potentially life-threatening disorders.6 The criterion standard for the evaluation of patients with suspected PE has evolved from pulmonary angiogram to ventilation–perfusion scan to computed tomography pulmonary angiography (CTPA) over the past four decades, but none represent a perfect diagnostic option. Even CTPA is only 83% sensitive and 96% specific, and it exposes patients to the risk of contrast dye-related allergic reactions, radiation, and nephrotoxicity, while likely identifying a number of clinically inconsequential subsegmental PEs.7-9 As a result of this combination of an imperfect test and the expectation of diagnostic certainty, CTPA ordering rates have increased while cause-specific PE mortality has remained static.10 In the United States, the overtesting phenomenon is manifest by PE prevalence rates that are consistently lower than in Canada, Europe, or Australia; U.S. clinicians evaluate lower-risk emergency department (ED) patients for PE.11 The workup for PE varies significantly from hospital to hospital.12 One pragmatic barrier to more consistent PE diagnostic approaches is the lack of high-quality evidence to facilitate Bayesian reasoning at the bedside. The Bayesian approach is a quantitative method using a combination of pretest probabilities and the likelihood ratios of diagnostic tests (elements of the history, physical examination findings, laboratory results, imaging studies, and clinical decision aids) to sequentially refine the posttest probability of a particular disease using deductive reasoning.13 The most frequent criticism of the Bayesian approach is that evidence-based estimates of pretest probabilities are lacking.14 It is for this reason that the Academic Emergency Medicine Evidence-Based Diagnostics series includes an improved understanding of specific pretest probabilities among its objectives.15 The diagnosis of PE is made even more complex during pregnancy, since serum D-dimer levels are often elevated, the natural physiology of pregnancy can mimic classic signs and symptoms of venous thromboembolism (including edema and dyspnea), and two patients simultaneously share the risks of radiation.16 Traditional medical teaching typically classifies pregnancy and the peripartum period as PE risk factors, despite the fact that existing PE clinical decision rules (e.g., the Wells criteria and the Geneva score) do not include pregnancy as a risk factor. In this issue of Academic Emergency Medicine, Kline et al.11 definitively disprove the myth of pregnancy as a risk factor for PE via a meta-analysis of 17 studies of 25,339 patients, including 506 pregnant patients. Their meta-analysis demonstrates minimal heterogeneity and a prevalence (pretest probability) of PE of 4.1% in pregnant patients compared with 12.4% in nonpregnant patients. Readers should be cautious not to extrapolate these results to the postpartum patient, because that subset of patients is actually at increased risk of PE.17 How does this article improve our ability to appropriately evaluate pregnant ED patients with potential PE? First, it provides disease-specific, high-quality, international data upon which to confidently base estimates of pretest probability for one of the most challenging diagnoses. Test and treatment thresholds for PE in pregnancy can now be estimated to maximize patient benefit and minimize iatrogenic harm.18, 19 In conjunction with validated PE clinical decision aids or physician gestalt, physicians can risk-stratify a substantial proportion of pregnant patients below the test threshold for either D-dimer or advanced imaging using this pretest probability. Kline et al. provide a rational and safe option to avoid overtesting and, as the science of PE diagnostics evolves, further advances that do not rely upon advanced imaging are likely to continue to emerge. For example, age-adjusted D-dimer values can now be used in many geriatric patients to exclude PE without further testing, thereby increasing the diagnostic efficiency of D-dimer testing.20 A pregnancy-adjusted D-dimer may very well do the same for peripartum patients. Second, these data provide a tangible example of a medical myth passed from one generation of physicians to the next due not to lack of data, but to the lack of systematic reviews to reveal an overarching perspective of the entirety of medical literature for a topic. This unfortunate reality is true for both therapy and diagnostic/prognostic research, but the Evidence-Based Diagnostics series provides a resource for physicians, educators, researchers, guideline developers, and policy-makers to find, publish, and reference definitive ED-specific diagnostic science. The next stage of knowledge translation to reduce unnecessary testing for PE in pregnancy likely includes decision support with clinician feedback, which has been found to decrease the use and increase the yield of ED CTPA, along with the dissemination of the risks of medical radiation and acceptable testing thresholds.21-23 While the provision of bedside physicians with pretest probabilities has been shown to reduce radiation exposure and the cost of care in low-risk patients,24 EM now needs to replicate these reductions in PE test ordering in nonresearch settings. The first “Preventing Overdiagnosis” conference, at which both of us (along with our colleague Jeremiah Schuur) spoke about overtesting of PE in the ED, established a hierarchy of short-term objectives for researchers, educators, policy-makers, and communication groups.5 Defining clinical decision points via systematic reviews was a top priority for the researchers, while developing decision support resources was a top priority for the educators. The Evidence-Based Diagnostics series and the manuscript by Kline et al.11 both meet these objectives for our specialty of EM. Two additional conferences have also been developed to address these issues. The next two Academic Emergency Medicine consensus conferences will be “Diagnostic Imaging in the Emergency Department: A Research Agenda to Optimize Utilization” in 2015 in San Diego and “Shared Decision-Making in the Emergency Department: Development of a Policy-Relevant, Patient-Centered Research Agenda” in New Orleans in 2016. PE is one prominent EM issue for which overimaging occurs, but the 2015 consensus conference will certainly identify others, including acute coronary syndrome, abdominal pain, and mild blunt trauma. Each of these conditions has an expanding body of research with which to evaluate the safety and efficiency of more focused advanced imaging. The 2016 consensus conference will provide a research agenda for ethical, pragmatic, meaningful shared decision-making between patients/families, emergency physicians, and consultants, while focusing efforts onto providing the best care for individual patients in the chaotic ED milieu. In hypothetical settings, low-risk patients often defer further testing when informed of the risk of testing.25, 26 Access to high-quality, ED-specific diagnostic evidence including pretest probabilities, likelihood ratios, and test-treatment thresholds will continue to be essential for physicians having conversations with patients regarding risks and benefits. Without this evidence, we will simply continue to pass our EM myths—like that of the increased risk of PE during pregnancy—down from generation to generation.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.011 | 0.055 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.002 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.002 | 0.001 |
| Research integrity | 0.003 | 0.018 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; both teacher heads agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".