On “Effectiveness of trigger point dry needling…” Cotchett MP, Munteanu SE, Landorf KB. Phys Ther. 2014;94:1083–1094.
Bibliographic record
Abstract
The study by Cotchett et al1 evaluated the effectiveness of dry needling on patients with plantar heel pain in the context of a parallel-group randomized controlled trial. The study was carefully planned and executed. Cotchett et al reported a very attractive number needed to treat of 4 and a statistically significant difference in mean scores in favor of the dry needling group over the sham treatment group. However, the authors noted the between-group difference was lower than their declared value for the minimal clinically important difference. In making this statement, it appears that Cotchett et al have equated the magnitude of an important within-patient change with that of an important between-group difference. Although what constitutes a clinically important difference is clearly a personal decision, previous work has shown that an important between-group difference will almost always be substantially smaller than an important within-patient change. Goldsmith et al,2 for example, found that an important intervention-control group difference was approximately 50% of an important within-patient change. Cotchett et al applied an estimate for the minimal important between-group difference based on a previous study by Landorf et al,3 who applied the estimation method of Jaeschke et al.4 Jaeschke et al defined the minimal important difference (MID) as “the smallest difference in score in the domain of interest which patients perceive as beneficial and which would mandate, in the absence of troublesome side effect and excessive cost, a change in the patient's management.”4 Their estimation method applies a 15-point global rating of change (descriptors range from “a very great deal worse” to “a very great deal better”) as the reference standard for change. The MID for the outcome measure of interest is quantified as the mean difference between patients reporting no change and patients reporting a small change on the reference standard. Although it may represent a reasonable estimate of a within-patient change, it will overestimate the size of an important between-group difference. The reason for this is that the estimate is based on the difference between one group where, according to the reference standard, 100% of patients have changed an important amount, and a second group where 100% of the patients were categorized as having no change. In real-world settings, not all patients in the intervention group will achieve an important improvement, and not all patients in the control or sham treatment group will remain unchanged. The effect of this within-group heterogeneity in change will be to pull the mean scores of the intervention and control groups closer together. The extent to which the means approach each other is dependent on the proportion of patients in the control and intervention groups who achieve an important change. To assist in the interpretation of the results of the study by Cotchett et al, it would be helpful to have the specific proportion of successes (ie, those meeting the criteria for a within-patient MID) in the dry needling and sham treatment groups and the difference in proportions of these groups with 95% confidence interval on the difference.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.022 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.002 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.002 | 0.002 |
| Scholarly communication | 0.002 | 0.002 |
| Open science | 0.002 | 0.000 |
| Research integrity | 0.020 | 0.017 |
| Insufficient payload (model declined to judge) | 0.005 | 0.006 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".