A Commentary on “Radiofrequency ablation reduces pain for knee osteoarthritis: A meta-analysis of randomized controlled trials” (Int J Surg 2021; 91:105951)
Bibliographic record
Abstract
Dear Editor, We took great interest to read the article by Li et al. entitled “Radiofrequency ablation reduces pain for knee osteoarthritis: a meta-analysis of randomized controlled trials” [1] published online in July 2021 in the International Journal of Surgery. The article clearly stated that radiofrequency ablation (RFA) was more effective in relieving pain and promoting functional recovery in patients with knee osteoarthritis. While we applauded the authors for their work on this topic, some methodological errors captured our attention on clinical context unrelated findings that might lead to biased results. We would like to have the following comments on this study. First, and most importantly, the authors specified multiple outcomes without clearly identifying primary and secondary outcomes. Most current studies typically only select an optimal indicator as the primary outcome due to statistical complexity of co-primary outcomes [2]. Furthermore, when all the outcomes were treated as co-primary outcomes in this study, the threshold for type I error rate should then be adjusted for each of the co-primary outcomes to 0.05/n (2-sided) according to the Bonferroni-Holm correction to maintain an overall familywise error rate of less than 0.05 [3]. More importantly, the threshold of statistical significance for pain scores should also be adjusted for repeated measurement data (i.e., 99% CI, P < 0.01) [3,4]. Second, the failure of the authors to define minimal clinically important difference (MCID) is concerning to us because MCID is critical in interpretation of clinical context. For pain intensity score, an 1- point decrease has been considered as MCID [4]. For the Western Ontario and McMaster Universities Arthritis (WOMAC) index, clinical efficacy of treatment was defined by a MCID decrease of 15-points [5]. Apparently, none of the results in this current meta-analysis had any clinically relevant benefits in the clinical context. Outcome measures with defined MCID are important to allow clinicians, researchers, and patients to understand changes in a manner that is clinically relevant. Altogether, the results coming from this meta-analysis have the potential to produce incorrect and misleading results. Third, six databases were searched with no language restriction in this study. However, the results could be more convincing if the authors were to search the databases more exhaustively to include Wanfang Data, Scopus, China National Knowledge Infrastructure (CNKI), BIOSIS preview, NLM Gateway, and ClinicalTrials.gov. Besides, a manual search strategy should also be included in this study. Finally, there are several minor shortcomings in this study which we would like to point out: (a) as the efficacy of RFA depends on the types of RFA, why no subgroup analysis was performed on this factor? Moreover, with high heterogeneities in the results, a sensitivity analysis should also be conducted in this meta-analysis. These heterogeneities can affect interpretation of results. We would also like to suggest to use a multivariate meta-regression model to detect whether there are any heterogeneities in the subgroups, rather than relying on a single subgroup-analysis or sensitivity analysis [4]; (b) It would be better for the authors to present their results with trial sequential analyses and to include data to determine whether the evidence was reliable and conclusive; (c) the authors did not sufficiently follow the PICOS format (Participants, Intervention, Comparison, Outcomes, and Study Design) for the inclusion and exclusion criteria according to the PRISMA guidelines; (d) funnel plots are essential to determine whether there is any publication bias in a study. These were not provided in this study. To us, funnel plots should visually be inspected, and the Egger’s linear regression test be performed to assess publication bias if at least ten trials were identified; (e) the authors did not explicitly specify measurements of pain intensity by using a numeric rating scale or a visual analogue scale; (f) finally, all the pooled results were lacking in specific units (i.e., millimeters or centimeters), which make reading of the already complex manuscript even more difficult. We appreciate that Li et al. have provided us with an important meta-analysis which can be used as a guide for clinical decision-making. However, as this meta-analysis drew conclusions based on unreliable statistical methodologies, correction of the above stated faults may result in different conclusions that can be drawn from the present meta-analysis. Ethical approval Not Applicable. Sources of funding None. Author contribution Yi-Feng Ren, Liu-Yi Long and Jing Dong conceived, designed, and planed the study. Liu-Yi Long supervised the study. Jing Dong and Yi-Feng Ren drafted the manuscript. Yi-Feng Ren and Liu-Yi Long critically revised the manuscript for important intellectual content. Research registration unique identifying number (UIN) Not Applicable. Guarantor Yi-Feng Ren and Liu-Yi Long. Provenance and peer review Commentary, internally reviewed. Declaration of competing interest None.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.026 | 0.009 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.016 | 0.026 |
| Bibliometrics | 0.004 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.006 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; both teacher heads agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".