In Reply: Misleading Article by Goertz et al
Bibliographic record
Abstract
Thank you for the opportunity to reply. The authors accuse us of misleading readers; below are the reasons why we respectfully disagree. Large number of articles screened is irrelevant: What is relevant is the robustness of the evidence contained in the 6 studies reviewed and their ability to support any reasonable conclusions or policy implications. Alternate interpretation: We stated it was our opinion1Goertz CM Hurwitz EL Murphy BA et al.Extrapolating beyond the evidence in a systematic review of spinal manipulation for non-musculoskeletal disorders: a fall from the summit.J Manipulative Physiol Ther. 2021; 44: 271-279Abstract Full Text Full Text PDF PubMed Scopus (18) Google Scholar that the sweeping interpretation presented in the Summit article was not supported by the 6 studies cited given the paucity and variability of the studies, small sample sizes, and lack of documented biological plausibility connecting the conditions.2Côté P Hartvigsen J Axén I Leboeuf-Yde C Corso M Shearer H et al.The Global Summit on the efficacy and effectiveness off spinal manipulative therapy for the prevention and treatment off non-musculoskeletal disorders: a systematic review of the literature.Chiropr Man Therap. 2021; 29: 8Crossref PubMed Scopus (13) Google Scholar Implied rationale: We are not sure what an ”implied rationale” is when applied to hypothesis-generated research. Regardless, only 1 study offered “correction of subluxation” as a putative causal mechanism while the others referenced the autonomic nervous system or influence on “various central descending inhibitory pathways,” or failed to articulate a biological rationale. Thus, we disagree with the assertion that all 6 studies had the same implied rationale. Statistical significance: The authors state that statistical significance was considered in connection with clinical significance in the Summit process. However, this was not true for trials with a P value of >.05. This concept is particularly important, as none of the accepted trials were large enough to be definitively negative or rule out clinically meaningful effects. As Rothman states, “It is easy to declare that a result is not statistically significant, falsely implying that there is no indication of an association, rather than to consider quantitatively the range of associations that the data actually support.”3Rothman KJ. Six persistent research misconceptions.J Gen Intern Med. 2014; 29: 1060-1064Crossref PubMed Scopus (227) Google Scholar Concern for scientific rigor for policy: Our article focused on concerns regarding the lack of rigor used to arrive at Summit article conclusions and policy implications.1Goertz CM Hurwitz EL Murphy BA et al.Extrapolating beyond the evidence in a systematic review of spinal manipulation for non-musculoskeletal disorders: a fall from the summit.J Manipulative Physiol Ther. 2021; 44: 271-279Abstract Full Text Full Text PDF PubMed Scopus (18) Google Scholar In fact, existing methodological standards exist for the rigorous translation of evidence into policy, such as the GRADE Evidence to Decision Frameworks.4Alonso-Coello P Schünemann HJ Moberg J et al.GRADE Evidence to Decision (EtD) frameworks.BMJ. 2016; : 353Google Scholar Unfortunately, such methodologies were not part of the Summit process. We did not question the scientific rigor of the systematic review process outlined in the Côté article.2Côté P Hartvigsen J Axén I Leboeuf-Yde C Corso M Shearer H et al.The Global Summit on the efficacy and effectiveness off spinal manipulative therapy for the prevention and treatment off non-musculoskeletal disorders: a systematic review of the literature.Chiropr Man Therap. 2021; 29: 8Crossref PubMed Scopus (13) Google Scholar Concern for best available evidence to inform policy: Given the paucity and variability of Summit studies, their small sample sizes, and the lack of documented biological plausibility connecting the 5 conditions included, they are not sufficiently robust to inform policy either individually or collectively. Clinicians and others should consider evidence: Each of us has dedicated our careers to the generation and dissemination of rigorous scientific evidence with the primary purpose of affecting clinical practice based on that evidence. In fact, it was our strong belief that policy must be evidenced-based and that clinicians and others should consider the strength of that evidence when formulating policy that formed the basis of our article.1Goertz CM Hurwitz EL Murphy BA et al.Extrapolating beyond the evidence in a systematic review of spinal manipulation for non-musculoskeletal disorders: a fall from the summit.J Manipulative Physiol Ther. 2021; 44: 271-279Abstract Full Text Full Text PDF PubMed Scopus (18) Google Scholar We disagree that we have misled anyone. Rather, we have elucidated the ways in which the current available evidence does not meet the level required for the rigorous development of policy implications. Extrapolating Beyond the Data in a Systematic Review of Spinal Manipulation for Nonmusculoskeletal Disorders: A Fall From the SummitJournal of Manipulative & Physiological TherapeuticsVol. 44Issue 4PreviewThe purpose of this article is to discuss a literature review—a recent systematic review of nonmusculoskeletal disorders—that demonstrates the potential for faulty conclusions and misguided policy implications, and to offer an alternate interpretation of the data using present models and criteria. Full-Text PDF
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.011 | 0.142 |
| Meta-epidemiology (narrow) | 0.002 | 0.002 |
| Meta-epidemiology (broad) | 0.004 | 0.003 |
| Bibliometrics | 0.003 | 0.002 |
| Science and technology studies | 0.004 | 0.005 |
| Scholarly communication | 0.008 | 0.009 |
| Open science | 0.005 | 0.005 |
| Research integrity | 0.051 | 0.061 |
| Insufficient payload (model declined to judge) | 0.011 | 0.016 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".