Welcome to the Real World: Do the Conditions of <scp>FDA</scp> Approval Devalue High‐sensitivity Troponin?
Bibliographic record
Abstract
In January 2017 the Food and Drug Administration (FDA) approved the first high-sensitivity cardiac troponin (hs-cTn) assay for use in the United States: the fifth generation hs-cTnT assay manufactured by Roche Diagnostics. This landmark decision finally enables Americans to benefit from the same improvements in diagnostic technology that the rest of the world has utilized for some 6 years. Improved analytical sensitivity and precision of hs-cTn assays allows accurate detection of small concentrations of cTn, which in turn allows earlier detection of cTn rise in patients with acute myocardial infarction (AMI). This means that the time period between serial samples can be shortened to as little as 1 hour.1 Even better for our crowded EDs, the ability to detect very low cardiac troponin concentrations means that we may be able to "rule out" AMI with just a single blood test. In fact, several prospective studies have recently demonstrated that one very low hs-cTn measurement (below the limit of detection) at the time of ED arrival has high sensitivity and negative predictive value (NPV) for AMI.2-4 However, due to concerns about imprecision at low concentrations, the FDA does not allow hs-cTnT concentrations to be reported below 6 ng/L. This cut point is above the assay's limit of detection (LoD) and limit of blank (LoB: 2 SD above the mean of results obtained when testing a sample with no troponin), which are 5 and 3 ng/L, respectively. This potentially significant decision by the FDA begs an important question: can a single measurement of hs-cTnT rule out strategy be used in the United States? In this issue, the data presented by McRae et al.5 begin to help us answer this question. Their study presents a real world evaluation of hs-cTnT use at five Canadian centers, including a total of over 7,000 patients. This retrospective study used routinely collected laboratory data to estimate the diagnostic accuracy of a single test rule out protocol with cutoffs set at 3 ng/L (the LoB), 5 ng/L (the LoD), and 6 ng/L (the lower limit of the reportable range in the United States). From a purist academic perspective, we should not hide from the fact that this study (by its nature) has numerous limitations. As with all retrospective studies, there is a question about patient selection. Are these truly the patients we would consider using a single test rule-out strategy for in practice? There is also the important issue of verification bias. In a diagnostic study, all patients must be subjected to an adequate reference standard. For a study of this nature, that should include serial troponin sampling. However, less than half of the patients had a second sample drawn, which is likely to overestimate diagnostic accuracy. We should also raise questions about the adjudication process. Outcomes were only adjudicated in patients with initial hs-cTnT concentrations within the normal range for whom the database had recorded the presence of a major adverse cardiac event (MACE). This means that "false-positive" MACEs could be down-graded, but "false negatives" in the database could not be revised, again tending to overestimate diagnostic performance. Despite these methodologic limitations, this study presents very important evidence: 1) that many clinicians, in this health system, are already discharging ED patients with chest pain based in part on a single hs-cTnT measure; 2) that implementation of a single 6 ng/L hs-cTn strategy may not reduce health system resource utilization; and 3) a single 6 ng/L hs-cTn strategy is sensitive for detection of AMI, but lacks sufficient sensitivity for MACE at 30 days. While it can never replace robust evidence from prospective studies, this evidence from a "real-world" evaluation is extremely valuable. It tells us how the strategy might be expected to work when emergency physicians are "let loose" with it, without the rigid constraints of a trial protocol and far less potential for a Hawthorne effect. A study of this nature also enables data to be collected from a large patient cohort. Prospectively screening, consenting, and following up over 7,000 patients in the ED would be an immense undertaking, and even with a world-leading team of research assistants, is likely to take years. One clear takeaway from this study is that within this Canadian health system discharging patients after a single hs-cTnT measurement is a common occurrence. Over half of the cohort (52%; 3,691/7,130) was discharged with just one hs-cTnT measure. While their decision to discharge these patients was likely influenced by many other variables (history, ECG, risk factors, etc.), it is clear that many of their clinicians already believe that a single hs-cTnT measure is sufficient in a large subset of patients with acute chest pain. This data set also challenges the assumption that implementation of a protocolized single hs-cTnT strategy will decrease health system resource utilization and ED length of stay. A single 6 ng/L hs-cTnT rule-out strategy would identify patients with initial hs-cTnT measures below 6 ng/L for early discharge, while identifying patients with measures at or above 6 ng/L for further evaluation. At many health systems this type of protocol would be expected to enhance ED throughout and reduce overall resource utilization. However, implementation of this strategy within these seven Canadian hospitals may actually increase their resource utilization and ED length of stay. At baseline (using their current care patterns) only 3,439/7,130 (48%) patients remained in the ED/hospital for serial hs-cTnT sampling. However, the number of patients unable to be ruled out based on the 6 ng/L strategy was 4,121/7,130 (58%). Therefore, implementation would have identified an additional 682/7,130 (10%) patients for observation and serial sampling, prolonging length of stay, and increasing resource utilization. Most importantly the data suggest that, by itself, a single hs-cTnT test may be insufficient to rule out ACS using the 6 ng/L cutoff. While a single measure < 6 ng/L had very high sensitivity and NPV for a diagnosis of AMI, it did not perform as well at predicting MACE (all-cause mortality, AMI, and revascularization). The sensitivity of a single hs-cTnT measure < 6 ng/L for 30-day MACE was 95% with a lower bound of the 95% CI of 93%. While the importance of missing MACE outcomes (particularly revascularizations) is a matter of intense debate, missing one in 20 of those who develop a MACE is unlikely to be acceptable to patients and clinicians in the United States given the important prognostic and medicolegal consequences. By comparison the use of hs-cTnT at LoD and LoB produced sensitivities for 30-day MACE of 97 and 99.5%, respectively, suggesting that the conditions of FDA approval may be decreasing the utility of hs-cTn assays in the United States. Ultimately, this highlights the importance of accounting for other clinical information, such as the ECG and a patient's symptoms. If we are to achieve "single-test rule out" with hs-cTnT in the United States, under the current FDA conditions, it will be imperative to evaluate hs-cTnT alongside clinical decision rules and risk scores. Now that hs-cTn assays are becoming commercially available in the United States we look forward to robust prospective evaluations of hs-cTn strategies in U.S. ED patients with acute chest pain designed to determine the most safe and effective way to utilize this new technology.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.016 | 0.111 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.004 | 0.007 |
| Scholarly communication | 0.014 | 0.013 |
| Open science | 0.002 | 0.004 |
| Research integrity | 0.015 | 0.018 |
| Insufficient payload (model declined to judge) | 0.042 | 0.019 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".