Research that matters: setting guidelines for the use and reporting of statistics
Notice bibliographique
Résumé
The IEJ is currently developing a series of editorial guidelines in various areas of endodontics (De-Deus 2012, Zehnder 2012, Hülsmann 2013, Mannocci 2013, Peters 2013), with the aim being to enhance the quality of manuscripts and, of utmost importance, to ensure the discipline of Endodontology continues to improve. In the present editorial, a distinctive focus is given towards the establishment of guidelines for manuscripts that contain an element of statistical analysis. The accuracy and relevance of statistical analysis is a critical element of scientific reports and is able to influence the conclusions as well as the decision-making of clinicians (Krithikadatta & Valarmathi 2012). The medical and dental literature is replete with articles stressing the misuse of statistics and data misinterpretation (Altman 2005, Polychronopoulou et al. 2011, Lucena et al. 2012). Currently, the IEJ is enjoying a record level of submissions (Zehnder 2012), and it is no surprise to observe a number of submissions with inappropriate use of statistics. This misuse of statistics has recently been of concern to the editorial board of another journal (Fouad 2013), which highlighted the need to establish guidelines for the reporting of data and its statistical analysis. In the medical field, a group of editors (known as the Vancouver group) has met in a regular basis since 1978, to discuss and promote a series of Uniform Requirements for authors, including several statistical guidelines that are now followed by more than 500 medical journals (The Journal of the American Medical Association 1997). As the IEJ is a leading journal in the field of Endodontology, it should embrace these guidelines and proactively uphold its responsibility to improve the quality of endodontic research, by following this international trend. From a statistical viewpoint, several concerns are often apparent when critically appraising submissions or even reading published manuscripts. From sample size selection to data interpretation, various statistics-related errors can be identified. To provide a clear understanding of the problem, these statistics-related problems can be classified into three categories, according to the phase of the study: As a referee for this journal for the last 5 years, I have noticed that over 90% of the manuscripts submitted to me for refereeing contain some sort of inappropriate selection or misuse of statistics in at least one of these categories. In endodontics, it has been reported recently (among other statistically related problems) that 41% of 209 published papers used the wrong statistical procedure and 19% of the conclusions were changed after the application of the correct statistical approach (Lucena et al. 2012). Similar findings have been reported in other fields of dentistry (Polychronopoulou et al. 2011, Krithikadatta & Valarmathi 2012). This is likely to have a negative impact on the reliability of the current scientific information. I assume here that the readers of the IEJ are familiar with statistical terminology, because I must briefly demonstrate how those factors may impact on the results of the studies. The need to design a study alongside the statistical analysis is frequently overlooked by authors. Authors should fulfil basic requisites to guarantee the absence of bias that prevents a sound statistical evaluation. Common prerequisites usually missed by authors include as follows: The test of a hypothesis is the art of examining whether a variation between two (or more) sample distributions can be explained by chance or not. By statistical convention, the speculated hypothesis is wrong, and the so-called null hypothesis occurs by chance. The test will determine whether this hypothesis is right or wrong. That is why the null hypothesis receives its name, because it has to be either nullified (rejected) or not nullified (accepted) by the test. When it is nullified, it is possible to conclude that the data support the ‘alternative hypothesis’, which is the one originally speculated (Dancey & Reiddy 2007). Authors should present the null hypothesis that is to be accepted or rejected by the test of the hypothesis. Whatever the result of a test of an hypothesis, there is always the risk of sampling error, named as type I (α) and type II (β) errors (Chia 1997). Type I error is the incorrect rejection of a true null hypothesis, while type II is the failure to reject a false null hypothesis. In order words, they are, respectively, the false-positive and false-negative results. Both types of errors are problematic. A false positive in health sciences (finding disease where there is none) causes unnecessary worry or treatment, while a false negative (failing to identify disease where there is one) gives the unsafe misconception of health and prevents patients from seeking treatment. In non-clinical studies, α and β errors also have a negative impact, leading to treatment rejection (in case of false positive) or acceptance (in false negative) where it should be otherwise. The ideal way to avoid those errors would be to investigate all the population. Obviously, this is not a possibility and that raises the importance of a sample size calculation for any kind of hypothesis test. Larger sample sizes usually increase precision; however, an increase in sample size reaches a point where the effect over precision is meaningless. That is, there is an ideal sample size point, which varies according to the effect size and the power selected for the study. Another type statistical error with study design is the unplanned introduction of bias or confounding factors to the groups. Groups must be equal at baseline and display comparable conditions, by avoiding the prevalence of a given characteristic or specific feature to a specific group (Dancey & Reiddy 2007). This condition is generally met via sample randomization, where each sample has an equal chance of being allocated to any group. Various mechanisms can be used to accomplish randomization, but it works better for larger sample sizes. Lack of blinding is another sort of confounding that interferes with the reliability of statistics. Both subjects and examiners can be blinded to the nature of the test or the material/technique, avoiding them making unintentional additional efforts or decisions by knowing the group they are allocated or examining (Dancey & Reiddy 2007). A common flaw during data analysis is the inappropriate selection of the statistical test, which is largely the result of the failure to: The common way to obtain general conclusions from any data set is to assume the data have a certain distribution. The bell-shaped Gaussian distribution, also known as normal distribution, is by far the most used method (Dancey & Reiddy 2007). The Gaussian distribution is a mathematical ideal; however, few biological distributions, if any, really follow it. Larger sample sizes increase the probability of fulfilling the Gaussian distribution, leading to previous attempts to establish an upper sample size limit (from 30 to 100) from which Gaussian distribution is assumed (Erickson 1973, Saunders & Trapp 2004). Obviously, this cut-off point remains controversial and establishing one here is counterproductive. To evaluate non-Gaussian distributed data, alternatives are used, such as data transformation into Gaussian or the use of nonparametric tests, as they do not assume normality (Lucena et al. 2012). Unfortunately, this is not a straightforward decision, because nonparametric (rank-based) evaluations have less statistical power and are not advisable for small sample sizes (Dancey & Reiddy 2007). A hypothesis test is the action used to identify whether a controlled variable is significantly influencing the outcome variable under measurement. In statistics, those variables are, respectively, defined as independent and dependent. The clear identification of those variables is critical to the selection of the appropriate statistical test. For instance, if one is studying hard-tissue accumulation after the use of three different rotary systems, clearly there is a single independent variable (rotary techniques) and a single dependent variable (accumulation of hard debris). A one-way anova (provided a Gaussian distribution is present) is certainly the best choice. However, if in the same set-up, the authors decide to also vary the irrigation procedure (syringe only or ultrasonic) for all rotary groups, it introduces an additional independent variable (irrigation), which implies the need for a two-way anova procedure. This test provides three respective P-values for the effects of (i) the instrumentation technique, (ii) the irrigation procedure and (iii) the interaction of those two independent variables. It is wrong to provide a one-way procedure to instrumentation-derived data and another one to irrigation data, because there is a real chance of increasing type I errors. It is also not advisable to create groups according to the interaction of the two independent variables (e.g. group 1 – instrumentation A + syringe irrigation, group 2 – instrumentation A + ultrasonic, etc.) and perform a one-way test, because a single P-value is generated, loosing information on the isolated effect of instrumentation and irrigation over hard-tissue accumulation. Precise application of hypothesis tests according to the number of independent variables was made, for example, by Santos et al. (2010) and Maalouf et al. (2013). As for this third category of statistically related problem, it is critical to: An overestimation of the P-value interpretation (Dancey & Reiddy 2007, Polychronopoulou et al. 2011, Lucena et al. 2012) is also of particular relevance. P-value interpretation is considered by many as a sound way to test a hypothesis. It is defined as the probability to which an observed difference among events is likely to be true in the population or is a result of random choice of samples (i.e. occurs by chance). As a standard, authors should set the significance level (α) that is considered to be a landmark for decision-making on accepting or rejecting the null hypothesis. It may vary from 1% to 10%, but usually is adopted at a 5% level (0.05), which is still quite arbitrary. The P-value is a unit-free value that serves only to drive authors to accept or reject the hypothesis based on the α level previously established. However, it can be misleading when reporting only P < or > 0.05, for example a 0.051 P-value is considered as non-significant, whereas a 0.049 is significant. Thus, reporting the exact P-value is more useful and allows readers to decide for themselves whether to consider the results significant. Another common misinterpretation is to consider the P-value as an indication of the size of a given effect. For instance, comparing a P = 0.01 with a P = 0.05 does not mean that a given experimental effect is stronger than the control, based on the lower P-value. It just says that only one sample from a population of 100 is likely to display a difference by chance when P = 0.01, and 5 out of 100 when P = 0.05. Certainly, the first option reduces the probability of chance and increases the certainty about a true difference between the groups, but it says nothing regarding the size of the effect. As the P-value does not indicate the size of the effect, it brings us to another statistically related problem: a perceived overestimation of P-value interpretation (Dancey & Reiddy 2007, Polychronopoulou et al. 2011). The straightforward consequence of over reliance on P-values is that its interpretation may be of little or no clinical significance. For instance, while comparing treatment A and B regarding the time required to heal an endodontic lesion, a significant 10-day difference between treatments is of limited clinical relevance. However, if the results are only P-value-based and not interpreted further, it is likely to influence the clinicians’ decision-making with a potential preference for one treatment, even though the effect size is low (only 10 days). This could lead a given treatment to be abandoned prematurely or to suffer from a reduction in interest. In such cases, the inclusion of confidence intervals (CIs) is critical to highlight the range of variation that occurs to 95% of the samples (Dancey & Reiddy 2007, Polychronopoulou et al. 2011, Lucena et al. 2012). As the name suggests, the CIs indicate how confident one can be in the observed results, by providing the range within which the differences are likely to be found (Altman 2005). As CIs use the units of the dependent variable, it allows one to conclude whether the interval of variance between the groups is in a range of clinical relevance, even when the null hypothesis has been rejected. Conversely, an intervention not found to be statistically significant might have a relevant clinical effect indicated by the CIs. In biology, provided a sample size of sufficient power has been selected, it is unusual that two given conditions are really equal. Thus, there is a real probability that a difference is found while comparing biological events (Dancey & Reiddy 2007). In endodontics, several variables are prone to influence authors and readers, limiting the results interpretation, as long as the P-value is the only evaluation performed. Some examples include leakage data, dentinal tubule sealer penetration, push-out strength, fracture resistance tests, etc. Even though it has been recognized for some time that there is a need to report CIs in addition to P-values, in the field of endodontic leakage, it is disappointing that only 1 publication of 90 contained CIs (Lucena et al. 2012). Thus, providing the P-value plus the CIs is considered as the far better and more complete way to report statistical findings. It should also be emphasized that reporting mean differences without a P-value is not acceptable, even if CIs are given. On the basis of this brief overview of the common statically related problems, the following guidelines are presented to authors: As a final consideration, and to reduce statistically biased manuscripts, it is suggested that raw data are submitted for every article accepted for publication. Raw data submission and archiving is already a policy of several high-impact factor journals. This allows Editorial Boards to recheck statistical analysis and provides an opportunity to advise authors to use more appropriate statistical procedures to prevent publication of unreliable data. The IEJ could also comply with the requisite of data archiving and sharing, making data accessible to scientific groups willing to reproduce the work or compare the results. Public access to raw files is recommended by some international organizations that regulate and provide funds for research in various countries. Moreover, the policy of requesting raw data could assist journals to reduce the not-infrequent movement towards data manipulation.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction distillée sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.
Scores Codex et Gemma par catégorie
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,006 | 0,051 |
| Méta-épidémiologie (sens strict) | 0,000 | 0,000 |
| Méta-épidémiologie (sens large) | 0,001 | 0,000 |
| Bibliométrie | 0,000 | 0,000 |
| Études des sciences et des technologies | 0,000 | 0,000 |
| Communication savante | 0,001 | 0,000 |
| Science ouverte | 0,001 | 0,000 |
| Intégrité de la recherche | 0,000 | 0,001 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,000 | 0,000 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule tête enseignante, pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».