MétaCan
Menu
Retour à la cohorte
Enregistrement W4306644116 · doi:10.1016/j.cjcpc.2022.10.003

Sounder Reporting of Study Results by Systematic Screening for Erroneous Interpretation of P Values and Statistical Tests in Cardiology

2022· editorial· en· W4306644116 sur OpenAlexaffabout
Frédéric Dallaire

Notice bibliographique

RevueCJC Pediatric and Congenital Heart Disease · 2022
Typeeditorial
Langueen
DomaineDecision Sciences
ThématiqueMeta-analysis and systematic reviews
Établissements canadiensCentre Hospitalier Universitaire de SherbrookeUniversité de Sherbrooke
Organismes subventionnairesnon disponible
Mots-clésInterpretation (philosophy)Medical physicsMedicineInternal medicineCardiologyStatisticsComputer scienceMathematics

Résumé

récupéré en direct d'OpenAlex

“This has nothing to do with pediatric and congenital cardiology.” Such was the response of a colleague when I showed him an earlier draft of this text. I disagreed. If we can improve the way we report cardiology research findings, then this has a lot to do with cardiology. I admit that the question I ask here applies to most of biomedical research: how can we steer away from the widespread misconceptions and erroneous interpretations of statistical tests and P values that are commonly seen in the biomedical scientific literature?1Altman D.G. Bland J.M. Absence of evidence is not evidence of absence.BMJ. 1995; 311: 485Crossref PubMed Scopus (1254) Google Scholar, 2Greenland S. Senn S.J. Rothman K.J. et al.Statistical tests, P values, confidence intervals, and power: a guide to misinterpretations.Eur J Epidemiol. 2016; 31: 337-350Crossref PubMed Scopus (1484) Google Scholar, 3Wasserstein R. Lazar N. The ASA statement on P-values: context, process, and purpose.Am Stat. 2016; 70: 129-133Crossref Scopus (3648) Google Scholar Our field is no exception: an informal survey of the research articles published in the first 2 issues of the CJC Pediatric and Congenital Heart Disease (CJCPC), as well as in 2 recent issues of CJC (volume 38, issues 6 and 8), revealed that >40% contained statements that may have been based on a misinterpretation of what a statistical test result really conveys (see Supplemental Tables S1 and S2). Given that CJCPC is still in its early start, why not set an example and work to better guide authors, reviewers, and editors on how to report the results of statistical tests. My goal here is not to offer a detailed exposé on biostatistics, but rather to give the readers some context on the scope of the problem, as well as to propose practical solutions that may very well be easy to put into practice. Nothing I will discuss here is new. It has been known for years that the scientific community has often oversimplified the results of statistical tests and erroneously used them as a marker of success of an experiment (P < 0.05 usually means “success”).4Nuzzo R. Scientific method: statistical errors.Nature. 2014; 506: 150-152Crossref PubMed Scopus (1064) Google Scholar Paths to correct the problem have also been proposed.3Wasserstein R. Lazar N. The ASA statement on P-values: context, process, and purpose.Am Stat. 2016; 70: 129-133Crossref Scopus (3648) Google Scholar Nevertheless, in a matter of a few months, several events made me realize that there is still a long way to go. First, I noticed that for manuscripts I was peer-reviewing, errors in the interpretation of statistical tests and inadequate reporting of effect size were widespread enough that I had been copy-pasting almost identical comments in many reviews. At the same time, I realized how often I was asked by peer reviewers and co-authors of my own scientific manuscripts to add statistical tests where it was inappropriate to do so. It seems others have had this experience as well.5Poole C. Low P-values or narrow confidence intervals: which are more durable?.Epidemiology. 2001; 12: 291-294Crossref PubMed Scopus (264) Google Scholar I then came across the American Statistical Association statement on P values, which has been around for a few years already.3Wasserstein R. Lazar N. The ASA statement on P-values: context, process, and purpose.Am Stat. 2016; 70: 129-133Crossref Scopus (3648) Google Scholar This statement opens by exposing an interesting and yet disconcerting circular phenomenon about why we still teach about the importance of P < 0.05 in graduate schools and medical schools: we teach it because many scientists and editors rely on it, and they rely on it because it was what they were taught.3Wasserstein R. Lazar N. The ASA statement on P-values: context, process, and purpose.Am Stat. 2016; 70: 129-133Crossref Scopus (3648) Google Scholar My own experience in teaching to first-year residents has taught me that a fair proportion have difficulty in teasing out what statistical tests can and cannot do. Nevertheless, I doubt that the root of the problems only lies in the teaching. These are complex concepts, and it is not surprising that they get oversimplified and somewhat distorted along the way. The American Statistical Association statement summarizes the important concepts very clearly.3Wasserstein R. Lazar N. The ASA statement on P-values: context, process, and purpose.Am Stat. 2016; 70: 129-133Crossref Scopus (3648) Google Scholar The result of a test statistic (the P value) only indicates how compatible the data are in regard to a statistical model. When this statistical model is used to test the null hypothesis, which applies to most cases in biomedical science, the P value becomes a measure of how much compatibility there is between the observed data (the results of the study) and the hypothesis that there is in fact no effect or no difference between groups (the so-called null hypothesis). When P values are small, it suggests that the data observed are incompatible with the hypothesis that there is no effect. It does not—and cannot—say more than that. Very importantly, the P value does not tell the investigator whether the treatment actually works, or if there is a difference between groups (the so-called alternative hypothesis). It does not say anything about the strength of an association or about the size of an effect, and it certainly cannot be used to state that there is no effect or no difference between groups. The common mistake here is to take a P = 0.05 and conclude that there is only 5% chance that the results would be a false positive. To take a P value = 0.05 and deduce from it that there is a 95% probability of a true difference between groups is equivalent to saying that a lottery winner likely cheated because the probability of winning by chance is low.3Wasserstein R. Lazar N. The ASA statement on P-values: context, process, and purpose.Am Stat. 2016; 70: 129-133Crossref Scopus (3648) Google Scholar,4Nuzzo R. Scientific method: statistical errors.Nature. 2014; 506: 150-152Crossref PubMed Scopus (1064) Google Scholar This common misconception that the P value may be used as a measure of the likelihood that the hypothesis being tested is true is especially sticky. Editors from very reputable journals have fallen for it6Pocock S.J. Stone G.W. The primary outcome is positive—is that good enough?.N Engl J Med. 2016; 375: 971-979Crossref PubMed Scopus (96) Google Scholar and persisted with this,7Pocock S.J. Stone G.W. The nature of the P value.N Engl J Med. 2016; 375: 2205-2206Crossref PubMed Scopus (4) Google Scholar even after being reminded that they were wrong.8Hu D. The nature of the P value.N Engl J Med. 2016; 375: 2205Crossref PubMed Scopus (4) Google Scholar This is a reminder that we should probably not wait passively until the problem goes away. We should instead be proactive and help authors and reviewers to screen for misconceptions and errors, and guide them towards solutions. I do not think that a full paradigm shift is needed here, and I do not envision that, in the short term, all scientific papers will suddenly fall back to Bayesian approaches to assess the likelihood of hypotheses. That said, an interesting step to take would be to provide authors and reviewers with a set of questions to pose before submitting or reviewing a manuscript. They are listed in Table 1, along with some examples of statements that should and should not be used. If the answer to any of these questions is “yes,” then there should be strong considerations to revise the manuscript.Table 1Questions authors should ask themselves before submitting a manuscript, examples of statements that should be avoided, and tips and suggestions on how to report resultsQuestions authors should ask before submitting a manuscriptCommon examples of statement that should be avoidedTips and examples of more accurate reporting of resultsAre P values used to determine whether the alternative hypothesis is true?Are P values used to measure the likelihood of a statement being true?“Our results support the hypothesis that there was an association between treatment A the outcome (P < 0.001).”“Our results do not support the hypothesis that there was an association between treatment A and the outcome (P > 0.05).”“The likelihood of an effect of treatment A on the outcome was high (P < 0.001).”“Our results favoured hypothesis A (P = 0.001), compared with hypothesis B (P = 0.04).”The probability that the alternative hypothesis is true cannot be measured by commonly used statistical tests. Rather than using P values, authors need to make a convincing scientific argument based on how strong the results are (effect size), on how likely the alternative hypothesis is in the context of the study (which statistical tests do not measure), and on the soundness of the study design.Is it implied that the presence or absence of an effect is based on a threshold of P value being reached?Are there P values in the text without mention of the effect size?Are there statements that are backed only by a P value?“Treatment A was associated with the outcome (RR = 2.3; P < 0.05), but treatment B was not (P > 0.05).”“Treatments A, B, and C were associated with the outcome. However, treatments D and E did not reach statistical significance (P > 0.05).”“We found an association between treatment A and the outcome (P = 0.01).”Statements on the presence of an effect or on the strength of an association need to be based on the effect size, not on the P value. The 95% CI may be used to provide additional information on statistical precision.“Treatment A was associated with the outcome (RR = 2.3, 95% CI: 1.8-2.7). The association was weaker for treatment B, and it was not considered clinically significant (RR = 1.2, 95% CI: 0.9-1.6).”“We found a clinically significant association between treatment A and the outcome, although the sample size was small and the 95% CI was wide (RR = 1.9, 95% CI: 0.4-5.3).”Is it implied that there is no difference between groups because a P value is high?Are P values used as a measure of the strength of an association, or as a measure of the size of an effect?“We found no association between treatment A and the outcome (RR = 1.3; P > 0.5).”“There was not difference in the median age between the treatment groups (P > 0.05).”“We observed a stronger association between treatment A and the outcome (P < 0.001), compared with treatment B (P = 0.03).”“There was a very strong association between treatment A and the outcome (P < 0.0001).”The appreciation of a difference between groups should be based on the effect size or on a measure of the magnitude of a difference, such as a standardized difference.10Austin P.C. Balance diagnostics for comparing the distribution of baseline covariates between treatment groups in propensity-score matched samples.Stat Med. 2009; 28: 3083-3107Crossref PubMed Scopus (3467) Google Scholar Large differences can have low precision because of the lack of an appropriate sample size. Very small differences may not be clinically significant yet be very precise because of a large sample size.“We could not find a clinically meaningful association between treatment A and the outcome, although the 95% CI was wide (RR = 1.15, 95% CI: 0.4-2.1).”“There was a difference between the mean age between group A and group B (12.2 vs 14.2 years, standardized mean difference: 0.8).”“The proportion of males at baseline was similar between groups (34% vs 38% in group A and group B, respectively, standardized mean difference: 0.12).”“We observed a stronger association between treatment A and the outcome (RR = 2.5, IC 95%: 2.1-2.9), compared with treatment B (RR = 1.5, IC 95%: 1.1-1.9).”“Compared with placebo, the effect of treatment A on the outcome was large, although the confidence interval was wide and overlapped with the absence of an effect (risk difference: 12.2%, IC 95%: −2.3% to 34.2%).”CI, confidence interval; RR, relative risk. Open table in a new tab CI, confidence interval; RR, relative risk. For those who may feel at a loss without the reassuring presence of a P < 0.05 as a marker of success, I suggest going back to the very interesting and practical tips offered by Charles Poole in 2001.5Poole C. Low P-values or narrow confidence intervals: which are more durable?.Epidemiology. 2001; 12: 291-294Crossref PubMed Scopus (264) Google Scholar In a nutshell, readers of scientific papers will get a better sense of the precision of the reported effect if confidence intervals are reported instead of P values.5Poole C. Low P-values or narrow confidence intervals: which are more durable?.Epidemiology. 2001; 12: 291-294Crossref PubMed Scopus (264) Google Scholar,9Rothman K.J. A show of confidence.N Engl J Med. 1978; 299: 1362-1363Crossref PubMed Scopus (240) Google Scholar These 3 numbers (the effect size and the 2 bounds of the confidence interval) enable the author to offer a nuanced interpretation of the importance of their finding (the effect size), as well as to comment on the level of precision afforded by sample size (the width of the confidence interval). The readers expect a sound interpretation of the size of the effect and its clinical implications, followed by comments on the precision of that effect size. A narrow confidence interval means that there is more precision, whereas a wide confidence interval calls for circumspection. The world being made of multiple shades of gray, authors should avoid turning it to black and white, and should thus be discouraged from falling back on a dichotomic and arbitrary mark of success and failure, such as P < 0.05, despite how tempting it may be. Similarly, authors should avoid turning the confidence interval into a simple dichotomic marker of success by simply looking if it crosses a certain value. Practically, I think a systematic screening of manuscripts before submission and during peer review is feasible. This may help authors and reviewers to avoid common pitfalls and raise awareness of the limits of the interpretation of statistical testing. Perhaps it may also help to shift the focus from one where reaching an arbitrary statistical mark is the gold standard, to one where we strive for sounder and more nuanced interpretation of study results, whatever they may be. This work complies with the Canadian Tri-Council Policy Statement on Ethical Conduct for Research Involving Humans. No funding was received for this study.

Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.

Comment cette classification a été obtenuedéplier

Prédiction distillée sur la base complète

Imitation des enseignants

Ni prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.

score de la tête « metaresearch » (Codex)0,050
score de la tête « metaresearch » (Gemma)0,262
Version: codex-gemma-dda1882f352aStatut de validation: machine_predicted_unvalidated
Catégories candidatesMétarecherche
Catégories consensuellesMétarecherche
DomaineSignal candidat: aucune · Signal consensuel: aucune
Devis d'étudeSignal candidat: Sans objet · Signal consensuel: aucune
GenreSignal candidat: Empirique · Signal consensuel: Empirique
Score de désaccord entre enseignants0,792
Score d'incertitude au seuil0,978

Scores Codex et Gemma par catégorie

CatégorieCodexGemma
Métarecherche0,0500,262
Méta-épidémiologie (sens strict)0,0000,000
Méta-épidémiologie (sens large)0,0060,001
Bibliométrie0,0000,001
Études des sciences et des technologies0,0000,000
Communication savante0,0000,000
Science ouverte0,0000,000
Intégrité de la recherche0,0000,000
Charge utile insuffisante (le modèle a refusé de juger)0,0000,000

Scores machine (provisoires)

Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.

Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.

Tête enseignante Opus0,249
Tête enseignante GPT0,470
Écart entre enseignants0,221 · la distance entre les deux têtes enseignantes sur ce seul travail
Statut de validationscore_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découle

Classification

machine, non validée

Prédiction automatique; les deux têtes enseignantes s’accordent sur ce qui est montré ici.

Devis d'étudeSans objet
Domainenon disponible
GenreEmpirique

Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».

En bref

Citations4
Publié2022
Routes d'admission2
Résumé présentoui

Explorer davantage

Même revueCJC Pediatric and Congenital Heart DiseaseMême sujetMeta-analysis and systematic reviewsTravaux en français237 207