Probability screening in manuscripts submitted to biomedical journals – an effective tool or a statistical quagmire?
Bibliographic record
Abstract
In recent years, ever-increasing examples of overt scientific misconduct, including plagiarism, duplicate publication, image manipulation, and data fabrication, have challenged traditional trust relationships between authors/clinical investigators, editors and readers of our medical journals. A number of major anaesthesia journals have been affected and, of greater importance, the veracity of the scientific record has been tarnished 1, 2. The editorial peer review process, that determines which articles are accepted for publication and how they are reported, is imperfect and often variable across journals. Nevertheless, within this framework, journal editors are duty bound to advance editorial policy, to plan innovative editorial content, and to ensure that published articles are novel and important, and that they are reported accurately and transparently. In the latter context, the editor's responsibility to address cases of suspected misconduct is often daunting owing to the impact of time and limited tools. Journals and their publishers generally follow the guidelines and recommendations of the flow diagrams offered by the Committee on Publication Ethics (COPE; see http://publicationethics.org). However, it is often difficult, and sometimes impossible, to verify the originality and accuracy of summary data reported in submitted articles. Journal editors now use plagiarism detection software programs such as CrossCheck® from iThenticate® (see www.ithenticate.com) to screen for plagiarism, and software programs such as Knowledge Finder® (www.kfinder.com) to assess the originality of articles submitted for publication. More recently, probability screening has emerged as another novel approach to test for the statistical validity of datasets. The purpose of this editorial is to consider the potential application of probability screening for datasets reported in manuscripts submitted to journals in consideration for publication. To highlight the challenge, during my term as Editor-in-Chief of the Canadian Journal of Anesthesia (CJA), we received an article in 2011 reporting the effects of a drug called colforsin daropate on diaphragmatic contractility in pentobarbital-anaesthetised dogs. The corresponding author was Dr Yoshitaka Fujii from Toho University School of Medicine, Japan. At the time of the article's submission, editors of the major anaesthesia journals had been reviewing ongoing concerns regarding the validity of Fujii's work, published over a number of years. An important Letter to the Editor published in Anesthesia and Analgesia in 2000 had raised credible concerns about the implausibility of the data on side-effects that were reported as almost always identical, in 47 articles written by Fujii, and published in five different anaesthesia journals 3. Regarding the article submitted to the CJA, there were several inconsistencies in the reported data; however, there was no clear method to assess the reliability of the dataset, even with independent statistical review. Concerns regarding the data were raised with the author and this eventually triggered an internal investigation at Toho University, as reported previously 4. Toho University subsequently determined that there was no ethical approval for the study and moreover, the data had been fabricated. Concurrently, in 2012, a seminal article written by Carlisle tested the data integrity of 168 randomised controlled trials written by Fujii 5. The article provided overwhelming statistical evidence that “the distribution of continuous and categorical variables reported in Fujii's papers, both animal and human, are extremely unlikely to have arisen by chance and if so, in many cases with likelihoods that are infinitesimally small”. After the issues came to light, and after five universities in Japan in which Fujii had worked previously were unable to vouch for the majority of his work, an unprecedented number of articles – 182 altogether – were subsequently marked for retraction by several journals (137 retracted so far, at time of writing). Fujii was subsequently dismissed from Toho University as a result of his egregious actions. In an ongoing initiative to refine further a reliable and sensitive statistical approach to detect potentially fraudulent data reported from randomised clinical trials (RCTs), Carlisle et al. take us one step further in applying established statistical methods for detecting ‘unlikely’ distributions of data in RCTs, in this issue of Anaesthesia 6. Their unique approach reflects a collaborative effort from editors of two key journals in our specialty: Anaesthesia and Anesthesia and Analgesia, and is based on the original method used to identify Fujii's fraudulent papers. However, while the conclusions from the 2012 article still stand, there was an identified small flaw in the original analysis 5. In the current paper, Carlisle et al. correct the chi-squared method reported previously, and compare its performance with analysis of variance (ANOVA) and another approach called Monte Carlo simulations, to test the probability of random sampling. Unsurprisingly, the mathematical analysis of the paper is complex. To reassure the readership, the article has been extensively reviewed by two experienced biostatisticians. In brief, the 2012 analysis of RCTs by Fujii showed that it was exceedingly unlikely that the baseline data of continuous response variables, including age, weight and height, could have been the result of random sampling. However, subsequent testing of the chi-squared method using simulation suggested inaccuracies resulting from imprecise means and the use of incorrect degrees of freedom and unmodified standard deviations. The article in this issue of Anaesthesia confirms that the corrected chi-squared method and the ANOVA method become inaccurate when the means are reported imprecisely. In contrast, fewer RCTs reported by Fujii had unlikely distributions when using Monte Carlo simulations compared with the other two methods, suggesting increased robustness with the Monte Carlo approach. The authors propose that: “the Monte Carlo analysis may be an appropriate screening tool to check for non-random (i.e. unreliable) data in randomised controlled trials submitted to journals”. To understand Carlisle et al.'s paper requires an appreciation of the fundamental importance of sampling, and random sampling in clinical trials. Sampling is concerned with selection of a subset within a population that will represent the characteristics of the entire population. There are several types of sampling, including – but not limited to – simple random sampling, systematic sampling, stratified sampling and cluster sampling. Randomisation is essentially a method to ensure an equal likelihood of subjects' being assigned to the intervention (treatment) group or control group in a clinical trial. The randomisation procedure is important because it: i) reduces the probability of selection bias if properly implemented; ii) facilitates blinding of treatment to investigators, participants and evaluators; and iii) permits the use of probability theory to test the likelihood that any difference in outcomes between intervention and control groups has occurred merely by chance 7. In fact, many clinical trials no longer use random sampling, instead choosing convenience samples from selected patient cohorts or hospitals, with strict cut-offs 8. While this approach may provide an element of practicality in trials, and greater precision of outcome variables of interest, the inherent risk is one of producing non-normally distributed data. This is a fundamental problem, as assumption of normality underlies most statistical tests. If this assumption is not satisfied, the resulting statistical tests become invalid. Use of Monte Carlo simulations for assessing non-random probabilities in RCTs, as reported by Carlisle et al., is a unique application of an established methodology. Briefly, Monte Carlo methods are a broad class of computational algorithms that rely on repeated random sampling to generate numerical results. The modern version was invented in the late 1940s by Stanislaw Ulam, while working on a nuclear weapons project that was subsequently instrumental in the Manhattan Project 9. His inspiration arose from playing solitaire – questioning the chances that a Canfield solitaire laid out with 52 cards will emerge successfully. In lieu of estimating by combinatorial calculations, a more practical method was to lay out the cards a number of times and simply observe and count the number of successful plays. The approach (under the code name: ‘Monte Carlo’) evolved into changing processes described by certain differential equations into a form that was interpretable as a succession of random operations. A Monte Carlo simulation uses repeated sampling to determine the properties of some phenomenon or behavior. At its basic form, the population of interest is simulated: “From the pseudo population, repeated random samples are drawn. The statistic under study is calculated in each pseudo sample, and its sampling distribution is examined for insights into its behavior” 10. The logic is easy to grasp; however, the execution is more challenging and the work is highly computer-intensive. In Carlisle et al.'s article, the simplest simulated scenario was a trial that reported the mean (SD) for two samples of a single baseline variable when each sample contained two measurements. Computers facilitated one million simulations of this scenario that were analysed for each of the four methods, generating 1 × 106 p values for each method. Results from repeated simulations showed that, of all the methods tested, Monte Carlo simulations were most reliable in testing for probability of random sampling in RCTs. There is a clearly identified need to screen for data in RCTs that may be inconsistent with random allocation, as experience has shown us over and over again, and for various reasons, that some authors will continue to cheat 11. Asking authors to submit (raw) original data may be useful; however, we recognise that authors who fake summary data could just as easily fake any corresponding raw data. As pointed out by Haldane, those who fake data would need to ensure that the data satisfy several layers of statistical cross-referencing, and three ‘orders of faking’, ensuring that: i) the mean values match what is expected; ii) the variance of means is within the range of those expected, and consistent with several inter-related variables; and iii) the results match the central limit theorem 12. From a practical perspective, Carlisle et al. report that Monte Carlo simulations with current computer technology can generate results fairly quickly: it required approximately one hour to simulate the data from the studies by Fujii, one million times each. Carlisle et al. suggest that researchers and journal editors might use parametric analyses for baseline variables in submitted report of clinical trials, and reserve Monte Carlo simulations to assess whether or not the sensitivity of the results is reasonable. However, it can take considerably longer to code the data and to verify that the codes functioned properly. Furthermore, the method may not be applicable to trials employing crossover designs or those reporting median values. In the latter instance, it would need to be established that means (SDs) estimated from median values could be reliably substituted into Monte Carlo simulations. Finally, the potential for bias in Monte Carlo simulations, due to the initial seed selection in the pseudo random number generators, is another potential limitation that must also be taken into consideration 13. At the present time, most journals editors would lack the necessary expertise to conduct these elegant simulations, and many of the 3500 + biomedical journals (why do we need so many?) lack the resources for an experienced biostatistician. In an ideal publishing world, all medical journals reporting original research would have an experienced biostatistician on their respective editorial boards, and authors everywhere would consult a biostatistician at trial design, as well as the analysis phase of each study. In reality, many editors are not adequately sensitised to the problem, and may be inclined to ‘look the other way’ for articles they have no intention of publishing. The majority of COPE's cases of suspected misconduct originate from a small number of journals, and it is exceedingly unlikely that these journals have all the problem manuscripts. As highlighted by the Fujii papers, there are instances when sharing of information amongst editors-in-chief may be crucial in exploring cases of suspected misconduct. In view the importance of confidentiality in the peer review process, editors, reviewers and authors should be aware that COPE has developed important guidelines to address this subject 14. I fully agree with Carlisle et al.'s assertion that “the analyses of published randomised controlled trials have little power to detect fraudulent data compared with internal auditing by university departments” 6. Proper vetting and internal peer review, for example by experienced authors and/or departmental/institutional leaders 15, might possibly increase the quality of material submitted to journals, and at the very least, provide an oversight process to ensure that reported research was actually conducted. This is a collective responsibility of all who are engaged in the research enterprise. In conclusion, Carlisle et al. are to be congratulated for their fine work. We anticipate and look forward to further updates that address some of the identified limitations of Monte Carlo simulations to assess the probability that reported means in RCTs reflect a truly random allocation of study subjects. In the meantime, editors of medical journals should follow the principles of the Code of Conduct and Best Practice Guidelines for Medical Journal Editors 16, and consider statistical consultation if applying Monte Carlo simulations as a screening tool for submitted manuscripts. I am former Editor-in-Chief, Canadian Journal of Anesthesia. No external funding declared.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.155 | 0.075 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.007 | 0.001 |
| Bibliometrics | 0.002 | 0.003 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.002 | 0.000 |
| Open science | 0.003 | 0.000 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.010 | 0.004 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; both teacher heads agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".