Conflicting results on the role of electrocardiogram in risk stratification for sudden cardiac death in childhood hypertrophic cardiomyopathy
Bibliographic record
Abstract
The research group at University College London introduced the concept of standardizing the quantification of risk by assessing risk of sudden death or equivalent malignant arrhythmia event occurring during the subsequent 5 years of follow-up after diagnosis. While the aims in the recent article by Norrish et al.1 published in this journal are laudable, it is imperative to highlight the methodological shortcomings underlying their analysis. The authors aim to assess the ability of the electrocardiogram (ECG) risk-score and other aspects of the 12-lead ECG to predict ‘major arrhythmic cardiac event’. They selected a subset (356/1029) of the multi-centre cohort of paediatric patients with hypertrophic cardiomyopathy (HCM) used to construct the HCM Risk-Kids algorithm, where ECGs had been archived. This subset however consisted mainly of recently recruited patients, and so median follow-up was only 3.9 years (interquartile range Q1–Q3 2.0–7.7).1 A major concern is that Norrish et al. did not restrict their analysis to include only those survivors who had reached 5 years of follow-up (n = 77 + 68 = 145), and the patients who had a ‘major arrhythmic cardiac event’ within 5 years of follow-up (n = 17). Instead they compared all patients without events, even if 25% of them had <2 years follow-up, with all patients who had major arrhythmic cardiac event, even if 8 of them had an event later than 5 years of follow-up. Thus, in fact, only 17 (patients with major arrhythmic cardiac event within 5 years) + 145 (patients with ≥5 years follow-up without event) = 162 of the patients are actually statistically informative about the end-point of major arrhythmic cardiac event within 5 years.1 This constitutes only 46% of the total ECG group analysed. The other 194 actually confuse the results as patients with high-risk ECG features and short follow-up might have had a major arrhythmic cardiac event later on within the 5-year span. Similarly, the ECGs of patients with major arrhythmic cardiac event after >5 years of follow-up might have evolved more ECG changes closer to the event as shown in other studies.2 Thus including these late major arrhythmic cardiac events will underestimate the sensitivity of the method. The poor statistical power of this study with short follow-up is illustrated by the finding that out of the five parameters found predictive in the HCM Risk-Kids algorithm, four (unexplained syncope, non-sustained ventricular tachycardia, left atrial diameter Z-score, and left ventricular outflow-tract gradient) fail to reach statistical significance in the subset.1 Studies with smaller number of included patients (n = 110–144), but with all included survivors having at least 5 years of follow-up, are perfectly able to reach statistical significance for risk factors.3,4 There are other methodological concerns. The ECG risk-score points should be expressed in whole numbers (0–14), and, as expected for a disease with a poly-genetic aetiology, have not been normally distributed in other studies of this score.2–5 Thus, it should not have been expressed with decimals, and should not have been represented by mean ± standard deviation, as done in the Norrish et al. study1 without the authors first demonstrating a normal distribution in their patients. To illustrate this point, we show the histograms of ECG risk-score distribution in the Swedish national cohort of patients with at least 5 years follow-up of all survivors3 (Figure 1). Frequency distribution histograms with superimposed expected normal distribution curves from patients included in Östman-Smith et al. study.3 From the top: the first electrocardiogram risk-score at diagnosis in the total cohort with mean follow-up of 13.4 year (106/110 had initial electrocardiograms available); middle panel: the first electrocardiogram-risk score from patients that at any point during subsequent follow-up suffered sudden cardiac death, re-suscitated cardiac arrest or appropriate internal cardiac defibrillator (ICD) intervention (major arrhythmic cardiac events); lowest panel: the first electrocardiogram risk-score in patients that suffered a major arrhythmic cardiac events within first 5 years of follow-up (n = 11). The different histograms confirm the non-normal distribution of the electrocardiogram risk-score both in the total cohort, and among those that have suffered a major arrhythmic cardiac events during follow-up, as well as the considerably better sensitivity of a cut-off of >5 points when only patients with major arrhythmic cardiac events within the first 5 years of follow-up are included. MACE, major arrhythmic cardiac events. Furthermore, out of the 1029 in the HCM Risk-Kids cohort, Norrish et al.1 state that 437 had no early ECGs and 92 had ‘poor quality traces’, leaving 500 patients, but only 356 were included, without reasons for further exclusions being explained in text. We note from the Figure 1A and B in Norrish et al.,1 that as typical of high-risk ECGs either the amplitudes of complexes from multiple leads overlap each other with standard magnification, or go outside the trace altogether.1 If those were the ECGs that were rejected as ‘poor quality’, there is a risk that high-risk ECGs might have been discarded. If used as illustrated this would cause difficulties in correctly calculating the ECG risk-score, which includes measurement of 12-lead ECG QRS amplitudes, and sometimes necessitates using ECG traces with reduced magnification. Fortunately, there is now another independent external validation of the ECG risk-score from Hospital Sick Children, Toronto,4 to compare with the Norrish et al. study results. This also contains tertiary centre patients like the HCM Risk-Kids cohort, but is more satisfactory from the methodological point of view, as all 144 eligible patients had at least 5 years of follow-up, and with 22 sudden cardiac death events within the first 5 years. Whereas Norrish et al. found a hazard ratio of 2.07 which did not reach statistical significance for an ECG risk-score >5 points, the Toronto results find that the C-statistic for a cut-off of >5 points was 0.76 for its ability to significantly discriminate between patients with and without sudden cardiac death events within 5 years.4 Thus, the Toronto results are similar to those from the Swedish National cohort where sensitivity was 97–100%,2,3 since Toronto data showed a non-normal distribution of risk-scores, a high sensitivity of 95% (Q1–Q3 77–100), positive predictive value of 28% (24–33), and negative predictive value of 99% (91–100).4 Median ECG risk-score for patients in Toronto with cardiac events was 8 (7–10),4 virtually identical to the Swedish data shown in Figure 1. The specificity was lower than in the Swedish cohort, 55% (46–65) vs. 73% (57–77),3,4 but it is perhaps not surprising that a geographical cohort will contain a larger proportion of low-risk patients than a tertiary centre: 20% of the Swedish cohort had an ECG risk-score of 0 points (Figure 1). In conclusion, there is no scientific basis for the inconclusive results obtained by Norrish et al.1 to deter other researchers from assessing what contribution ECG changes could make to improve risk-stratification for the electrophysiological event of a malignant arrhythmia. It would help this research field if authors studying risk factors for major arrhythmic cardiac event would standardize to assess prediction of major arrhythmic cardiac event during the first 5 years of follow-up, and only include survivors who have completed 5 years of follow-up. The authors were supported by grants from the Swedish Heart-and Lung Foundation (Number 20080510), and the Swedish state under the agreement between the Swedish government and the county councils, the ALF-agreement (ALFgbg-544981), Region Östergötland (ALF), the Strategic Research Area in Forensic Science, and FORSS (Medical Research Council of Southeast Sweden).
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.007 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".