Notice bibliographique
Résumé
The p-value statistic—often misused and misunderstood—has come under considerable criticism, but continues to be used in scientific papers Did you sleep through your statistics class in medical school? You probably weren't alone. Maybe that’s one reason so many excellent medical researchers feel at odds with statistical analysis. Or maybe the natural optimism of scientists leads them to report “statistically significant” results that later can’t be duplicated. Whatever the reason, one specific statistic—the p-value—is too often manipulated to become a sole determinant of significance, leading to results that can’t be replicated.1Ioannidis JPA Why most published research findings are false.PLoS Med. 2005; 2: 0696-0701Crossref Scopus (5866) Google Scholar It’s become one of the largest problems in biomedical research. It is not the method itself that is problematic, however, but the way it is used. According to Ronald Fisher, the “father” of the p-value, the method was only meant to inform the investigator whether to do more extensive studies, says Bruce Kaplan, MD, professor of medicine at the Mayo Clinic in Phoenix and professor of health solutions at Arizona State University. Dr. Kaplan adds that findings may be reproducible, but only if the researchers sample the same population. The p-value, the most widely used statistic in biomedical research, is defined as the probability of obtaining a result equal to or “more extreme” than what was actually observed when the null hypothesis is true. A small p-value, typically of less than or equal to 0.05, indicates strong evidence against the null hypothesis. “I don’t believe there is anything fundamentally wrong with p-values; the problem is the overreliance and degree of emphasis placed on p-values,” says Jesse Schold, PhD, director of outcomes research and medical informatics at Ohio’s Cleveland Clinic. Mark D. Stegall, MD, a clinical investigator and professor of surgery research at the Mayo Clinic in Rochester, Minn., agrees, commenting, “From a purely statistical perspective, there is nothing ‘wrong’ with p-values. The pitfall is that we put too much value in them. Our interpretation is the fault.” Use of the p-value is part of the medical culture, says Philip Halloran, MD, PhD, director of the Alberta Transplant Applied Genomics Centre at the University of Alberta in Edmonton. Dr. Halloran, the founding editor of the American Journal of Transplantation, says editors only see the successful tests done and not the entire raw data. He likens use of the p-value to a phenomenon called “looking for the pony.” It’s an old joke regarding a pile of manure and an optimistic child who claims “there must be a pony in there somewhere!” He adds, “If you start looking for anything at all in biomedical research, you’ll find something.” “The real problem isn’t the statistics; it’s the experimental design,” continues Dr. Halloran. “[Researchers] know they’re seeing things that can’t be reproduced. Editors are not being shown the complete design because someone has looked for the pony and found it, but [researchers] are reporting something from only one test, not the many that they’ve conducted.”Key Points•>Some experts say that the p-value is too often manipulated in biomedical research to become a sole determinant of significance, leading to results that can’t be replicated.•>Other statistics, such as confidence intervals, probabilities and power analyses, can also be used to complement findings.•>The American Statistical Association recently released a statement that outlines how to use the p-value, and encourages journals to stop using statistical significance to determine whether to accept an article. •>Some experts say that the p-value is too often manipulated in biomedical research to become a sole determinant of significance, leading to results that can’t be replicated.•>Other statistics, such as confidence intervals, probabilities and power analyses, can also be used to complement findings.•>The American Statistical Association recently released a statement that outlines how to use the p-value, and encourages journals to stop using statistical significance to determine whether to accept an article. David N. Ikle, PhD, principal statistical scientist at Rho, a research organization based in Chapel Hill, N.C., says, “There are at least two [statistical] issues of relevance to the transplant community. One is the potential for overinterpretation of highly significant effects estimated from very large transplant registry databases, when the effect size may be too small to be of clinical significance. Second, use of p-values is not necessarily problematic in the relatively small clinical studies typical in the transplant community since they do properly incorporate the sample size in the calculation based on the chosen statistical model. However, in those studies it is especially critical to employ principles of good study design, to carefully define robust clinical endpoints, to carefully manage the conduct of the study, and to follow the National Institutes of Health guidelines for ensuring rigor and reproducibility in biomedical research.” He adds, “In our work with dozens of researchers in the transplant community, we try to educate our collaborators in the potential pitfalls of overreliance on p-values and suggest alternatives where appropriate … it is still a challenge to overcome decades of research that relied solely on an arbitrary cutoff on p-values to determine the importance of research finding.” Dr. Schold notes, “There are other statistics, such as confidence intervals, probabilities, power analyses, that may also be used to complement findings. At the end of the day, research methods should be transparent.” “From a purely statistical perspective, there is nothing ‘wrong’ with p-values. The pitfall is that we put too much value in them. Our interpretation is the fault.” –Mark D. Stegall, MD In response to growing concerns about the p-value and the inability to replicate study findings, the American Statistical Association (ASA) released its first-ever statement on statistical significance and p-values in March 2016.2Wasserstein RL Lazar NA The ASA’s statement on p-values: Context, process, and purpose.Am Stat. 2016; 70: 129-133Crossref Scopus (3264) Google Scholar Ron Wasserstein, executive director of the ASA, notes that journals are the gatekeepers that can usher in a “post p ≤ 0.05” era. “If the [ASA] statement succeeds in its purpose, journals will stop using statistical significance to determine whether to accept an article. Instead, journals will accept papers based on clear and detailed description of the study design, execution and analysis … and they will be reported transparently and thoroughly enough to be rigorously scrutinized by others.” Whatever happens with biomedical literature in the future, it will certainly continue to include statistics. “Most of the Nobel Prizes that have been won in scientific fields have all been based on statistics and mathematics,” Dr. Kaplan says. Dr. Halloran says, “A better literature would take into consideration a more rigorous approach to experiment design that took bias into account. These researchers are genuinely honest, but, like all of us, they can be biased.”Medical Journals Sound Off on p-ValueAcademic journals vary in their requirements for statistical analysis from researchers. Following are a few examples:•>PLoS—The journal asks that authors quantify findings and present them with appropriate indicators of measurement error or uncertainty (such as confidence intervals), when possible. The journal asks that authors avoid relying solely on statistical hypothesis testing, such as p-values, which the journal guidelines say “fail to convey important information about effect size and precision of estimates.”•>JAMA—At the end of the case presentation, the journal requires the pertinent diagnostic test results and normal ranges be provided. Four plausible responses should be provided to answer the question “How do you interpret these test results?” There is no reference to p-values.•>NEJM—Statistical consultants review every paper and look for “appropriate” methods, ie, authors need to have used something appropriate to the data they are reporting and the type of study they are conducting. Additionally, NEJM does not require authors to submit raw data.•>The Lancet—The journal asks authors to provide p-numbers to four decimal points and requires that manuscripts adhere to international reporting guidelines. During peer review, papers are reviewed by at least three external experts and at least one independent statistician. The journal also does not request access to raw data. Academic journals vary in their requirements for statistical analysis from researchers. Following are a few examples:•>PLoS—The journal asks that authors quantify findings and present them with appropriate indicators of measurement error or uncertainty (such as confidence intervals), when possible. The journal asks that authors avoid relying solely on statistical hypothesis testing, such as p-values, which the journal guidelines say “fail to convey important information about effect size and precision of estimates.”•>JAMA—At the end of the case presentation, the journal requires the pertinent diagnostic test results and normal ranges be provided. Four plausible responses should be provided to answer the question “How do you interpret these test results?” There is no reference to p-values.•>NEJM—Statistical consultants review every paper and look for “appropriate” methods, ie, authors need to have used something appropriate to the data they are reporting and the type of study they are conducting. Additionally, NEJM does not require authors to submit raw data.•>The Lancet—The journal asks authors to provide p-numbers to four decimal points and requires that manuscripts adhere to international reporting guidelines. During peer review, papers are reviewed by at least three external experts and at least one independent statistician. The journal also does not request access to raw data. The National Institutes of Health (NIH) have issued new Scientific Rigor and Reproducibility criteria designed to “ensure robust and unbiased experimental design, methodology, analysis, interpretation and reporting of results.” Additionally, the NIH has developed four video modules with accompanying discussion materials that focus on integral components of reproducibility and rigor in the research endeavor, such as bias, blinding and exclusion criteria. The criteria and modules can be found at www.nih.gov/research-training. A program initiated at the University of California, Los Angeles (UCLA) in late 2014 and now offered nationwide through the National Kidney Registry as an “advanced donation program” allows medically and psychosocially acceptable donors to donate a kidney before their intended recipient receives a kidney. UCLA tells the story of 64-year-old Howard Broadman, who approached the institution with the idea of donating a kidney so that his 4-year-old grandson Quinn would be eligible to receive a kidney in the future. The donation took place in December 2014. “I know Quinn will eventually need a transplant, but by the time he’s ready, I’ll be too old to give him one,” Broadman told UCLA. “Why don’t I give a kidney to someone who needs it now, then get a voucher for my grandson to use when he needs a transplant in the future?” Under the umbrella of the National Kidney Registry’s program, nine other transplant centers in the United States have agreed to offer the voucher program.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,058 | 0,307 |
| Méta-épidémiologie (sens strict) | 0,003 | 0,002 |
| Méta-épidémiologie (sens large) | 0,004 | 0,002 |
| Bibliométrie | 0,012 | 0,008 |
| Études des sciences et des technologies | 0,004 | 0,027 |
| Communication savante | 0,023 | 0,018 |
| Science ouverte | 0,005 | 0,012 |
| Intégrité de la recherche | 0,011 | 0,017 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,057 | 0,040 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».