MétaCan
Menu
Back to cohort
Record W2916886576 · doi:10.1111/acem.13717

Prognosis Versus Diagnosis and Test Accuracy versus Risk Estimation: Exploring the Clinical Application of the <scp>HEART</scp> Score

2019· letter· en· W2916886576 on OpenAlexaff
Christopher Byrne, Cristian Toarta, Tim Holt

Bibliographic record

VenueAcademic Emergency Medicine · 2019
Typeletter
Languageen
FieldMedicine
TopicAcute Myocardial Infarction Research
Canadian institutionsUniversity of Toronto
FundersNational Institute for Health and Care Research
KeywordsMedicineTest (biology)Framingham Risk ScoreEstimationInternal medicine

Abstract

fetched live from OpenAlex

To the Editor: In a recent issue, Fernando et al.1 present a systematic review and meta-analysis assessing the prognostic accuracy of the HEART score for prediction of major adverse cardiac events (MACE) in adult patients presenting with chest pain at the emergency department (ED). The authors conclude that the HEART score has excellent performance for prediction of MACE (particularly mortality and myocardial infarction) in chest pain patients and should be the primary clinical decision instrument used for risk stratification of this patient population. While we agree with the utility of the HEART score, we take an alternate evidence-based medicine perspective on how to best demonstrate this utility to patients, clinicians, and other stakeholders. In this letter, we explore the concepts of prognosis, diagnosis, test accuracy, and risk estimation as they relate to the HEART score. Increasingly, clinicians and health researchers are recognizing the burden overdiagnosis can have on individual patients and health care systems. The usefulness of diagnostic decisions should be judged by whether patients classified with diagnosed disease do better than those classified without disease.2 This requires information about patient prognosis—in this case, the likelihood of future patient-centered outcomes in patients with a given HEART score. There is a case for prognosis to replace diagnosis as the framework for clinical decision making in patients presenting to the ED with chest pain. As the authors of this review acknowledge, there is a tendency for clinicians to overinvestigate chest pain patients, resulting in increased resource utilization without improved outcomes.3 In part, this is due to the lack of a good reference standard in the diagnosis of acute coronary syndrome (ACS). The diagnostic test accuracy literature for ACS in chest pain patients presenting to the ED most often uses clinical follow-up (e.g., future MACE) or a panel adjudicated final diagnosis based on clinical history and hospital course as the diagnostic reference standard.4 Both of these standards are flawed as they are subject to important incorporation, verification, and spectrum biases likely to misrepresent reported diagnostic test accuracy measures.5 As a result, ACS remains at its core a clinical diagnosis, with certain investigative tools (e.g., electrocardiogram, troponin) providing additional information supporting the diagnosis. At present, the Cochrane Collaboration has guidance on the performance of systematic reviews of diagnostic test accuracy,6 while a working group (Cochrane Prognosis Methods Group) attempts to reach consensus on how prognostic systematic reviews are best conducted.7 We highlight this because the authors of this review opted to perform a meta-analysis of diagnostic test accuracy rather than predictive ability. The rationale for this decision is that the HEART score is primarily used to “rule out” MACE in low-risk patients and so clinicians (and presumably patients) will be most interested in the accuracy of their screening decision. The authors state that when evaluating a decision instrument in the context of screening, the most important test characteristics are sensitivity, specificity, and likelihood ratios. We advocate that predictive values are the more clinically useful measures for this review's clinical question and should be routinely presented alongside other measures. Authors of prognostic reviews should also be challenged by peer reviewers to present their data in subgroups according to disease prevalence (e.g., low, intermediate, and high). The interested reader can then more reliably assess the performance of test characteristics across study populations with variable baseline risks of future MACE. Sensitivity and specificity are commonly taught to be fixed properties of a test that do not vary with disease prevalence. However, in clinical practice, the sensitivity and specificity of a test can vary with disease prevalence.8, 9 This phenomenon is known as spectrum or case-mix bias. Spectrum bias acknowledges the possibility that shifts in test performance may be at least in part due to case mix variation among study populations. Case mix variation can impact the sensitivity, specificity, and likelihood ratios of a given test. Among the studies included in this review, future MACE prevalence ranges from 1.1%10 to 29.4%.11 This is a wide prevalence range. It is reasonable to question the appropriateness of applying the reported pooled sensitivity (95.9%) for HEART score > 3 to such diverse populations. Assessing the external validity of the pooled sensitivity to variable populations is made more difficult by the fact that the review authors do not report a pooled prevalence of future MACE, though this can be calculated using data presented in Figure 2. To illustrate this concept further, the reported sensitivity in the low MACE prevalence study by Mahler et al.10 was 58%. Why is the sensitivity in this study so far from the pooled sensitivity of 95.9% reported in this review? The likely answer is spectrum bias. In fact, this American study involved patients in an observation unit who were already determined to be low risk by the treating clinician (on the basis of clinical assessment, non-diagnostic electrocardiogram and negative troponin).10 This population does not represent the chest pain patient we would typically be applying the HEART score to as observation units for low risk chest pain do not exist where we work. The study by Mahler et al.10 also presented predictive values, demonstrating that 0.6% (5/904) of patients with low-risk HEART scores suffered a future MACE. In this context, we believe that a negative predictive value of 99.4% is more helpful to the treating clinician in reaching a shared clinical decision than a sensitivity of 58%. When the HEART score is viewed as a test, there is also a risk that clinicians will attempt to “rule in” or “rule out” ACS using the HEART score. And can you blame them? After all, the population (ED patients with chest pain), index test (HEART score), reference standard (MACE), and target condition (“clinically significant” cardiac ischemia) alluded to in this review appear to provide the basic framework for an ACS diagnostic accuracy study. It is important to highlight that the HEART score was not designed to be applied to patients with a proven ACS. Patients with definite ACS at presentation are usually excluded from HEART score studies. These patients should be treated in the usual manner and referred for ongoing medical management and/or revascularization. The HEART score should be applied to those patients with chest pain, in which the diagnosis is uncertain, yet the clinician considers a cardiac etiology possible. Conceptually, we do not view the HEART score as a test. Many HEART score studies do not report measures of test accuracy, including those by authors involved with the HEART score's development.12 The HEART score is a set of clinical (history, risk factors, age) and investigative (electrocardiogram, troponin) criteria permitting estimation of a patient's future short-term risk of death, myocardial infarction, or revascularization. We liken it to the Wells’ score for pulmonary embolism, which is used to estimate a patient's short-term risk of pulmonary embolism. This pretest probability subsequently informs the decision to evaluate a patient with a D-dimer or diagnostic imaging.13 The clinical utility of the Wells’ score is not in ruling in or ruling out future thromboembolic events. Likewise, the clinical utility of the HEART score is not in ruling in or ruling out future MACE. Using terms like rule in or rule out encourages emergency clinicians to strive for perfection or no “misses” in the evaluation of these patients. These terms also perpetuate a view that for the HEART score to be valuable at the bedside, researchers must demonstrate the HEART score “outperforms” clinician gestalt. We need to accept that it is not possible to get ED patients with possible cardiac chest pain to a 30-day to 6-week MACE rate of zero, whether using the HEART score, clinician gestalt, or any other chest pain risk score. Rather than attempting to rule out future MACE, the clinician should ask the following question: is this patient's estimated risk modifiable by more observation and testing (e.g., serial troponin, cardiac stress test) over outpatient management of risk factors for coronary artery disease? We think Fernando and colleagues should be congratulated for their substantial work in this systematic review and meta-analysis of the HEART score. We appear to have the same intentions in mind but offer alternate views on how to evaluate and apply the HEART score literature.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.027
metaresearch head score (Gemma)0.238
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Theoretical or conceptual · Consensus signal: none
GenreCandidate signal: Commentary · Consensus signal: Commentary
Teacher disagreement score0.027
Threshold uncertainty score0.141

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0270.238
Meta-epidemiology (narrow)0.0010.001
Meta-epidemiology (broad)0.0030.002
Bibliometrics0.0020.002
Science and technology studies0.0010.002
Scholarly communication0.0040.003
Open science0.0030.001
Research integrity0.0090.012
Insufficient payload (model declined to judge)0.0050.002

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.205
GPT teacher head0.431
Teacher spread0.226 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designTheoretical or conceptual
Domainnot available
GenreCommentary

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations1
Published2019
Admission routes1
Has abstractyes

Explore more

Same venueAcademic Emergency MedicineSame topicAcute Myocardial Infarction ResearchFrench-language works237,207