Treatment for Chronic Sciatica: Do We Have New Evidence Showing That Discectomy Has Any Long-Term Benefits Over Conservative Care?
Bibliographic record
Abstract
Commentary The attempt to identify real advantages of surgical treatment over conservative treatment for lumbar herniation has spurred a continuing series of studies over the last 40 years. Despite varying in population, size, study design, treatment techniques, and patient condition, most studies have had very similar findings: that surgery provides faster pain relief, with improved scores, at earlier time points. However, in almost all studies, those differences decrease with time and vanish after 1 or 2 years. The article by Bailey et al. contains valuable information for surgeons considering prompt surgical intervention for chronic lumbar disc herniation rather than a more common nonoperative treatment regimen prior to possible surgery. In undertaking this randomized controlled trial (RCT) using intention-to-treat analysis, the authors took advantage of a feature of the Canadian health-care system with its inherent waiting periods of >6 months, which limited the ability of patients in the nonoperative group to cross over to surgery in the early periods of the study. Nevertheless, this study recorded 24 crossover events: 2 early events, 12 events between 6 and 12 months, and 10 events by 2 years. In addition, 8 patients in the surgery cohort did not have surgery. In the first 6 months, crossover events were indeed low. However, with the increasing number of crossover events after 6 months and then more after 12 months, the unique advantage in study design between this RCT, undertaken in Canada, and previously published RCTs appears to vanish. In contrast to several other prior studies mentioned by the authors, this study showed evidence of advantages of prompt discectomy rather than additional nonoperative treatment over the longer term of 2 years as well as at 1 year after treatment group assignment. Under close scrutiny, the differences at 1 year and earlier do appear to be meaningful, but the differences at 2 years are not substantial and do not refute the findings of most other related studies: that the difference between groups decreases over time, becoming clinically unimportant by 2 years. The interpretation of many published studies is made difficult by statistical methodology and metrics that are not readily familiar to most surgeons, but it is important to look deeper before accepting an author’s conclusions at face value. In their paper, the authors use and recognize a minimal clinically important difference (MCID) in the patient-reported outcome measures. Based on a study by Lauridsen et al., Bailey et al. chose 2 as the MCID for their leg pain score, meaning that any difference of <2 is recognized as not clinically meaningful1. However, the authors report the mean difference between treatment groups in their primary outcome measure, leg pain, to be 1.3, well below their acknowledged MCID, indicating that the difference is not clinically important even though it might be significantly different. The MCID can be determined in many ways with use of different methods, so perhaps the value of 2 is off-base2. Copay et al. compared several different ways of determining the MCID with use of similar data on similar outcome measures3. They found the best estimates to be 4.9 for the SF-36 physical component summary (PCS) score, 1.6 for leg pain, and 1.2 for back pain. Although the SF-36 PCS score reported by Bailey et al. is just above Copay’s proposed MCID (5.3 vs. 4.9) indicating clinical importance, the mean differences found by Bailey et al. for leg pain and back pain were both lower than Copay’s proposed MCIDs, indicating differences that are clinically unimportant. So what can we conclude? Although the results of the study by Bailey et al. offer strong evidence for a difference in outcome measures at 6 and 12 months as previously reported4, there was a very marginal or clinically unimportant difference between Bailey’s groups at 2 years. Additionally, readers should always be wary of the very real placebo effect in studies such as this, in which no blinding of patients or surgeons has been accomplished. Based on these considerations, the authors’ conclusion that “microdiscectomy is superior” at 2 years may be just a little bold. The choice to undertake surgery or nonoperative treatment prior to potential surgery should not be taken lightly, even for patients with chronic conditions, because surgery entails risks. The overall surgical complication rate at 2 years reported by Bailey et al. was 15% (with at least one surgery-related adverse event occurring in 12 of 80 patients who underwent surgery). The rate in a previously reported meta-analysis was similar, at 12.5% for open microdiscectomy5. An RCT design, as used by Bailey et al., selects patients blindly for assignment to one treatment or another and reports the results as means for each cohort. Physicians, however, should never blindly choose treatment based solely on what appears to be best for the average patient but should rely instead on a range of available information about individual patients, including pain level, function, mental state, length of symptoms, and above all, patient preference. It may be informative that, even in a socialized medical environment, 40 of the 64 patients initially assigned to the nonoperative treatment group did not go on to have surgery within the time frame of the study even though they could have easily done so at 6 months without cost. The guidance suggested by Legrand et al.6, that the best approach is to “let the patient choose between treatments,” remains valid. Those authors recommended that patients be informed that surgery does not modify the long-term outcome but can speed up recovery, at the expense of potential complications, most of which are reversible6.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.014 | 0.137 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.005 | 0.003 |
| Bibliometrics | 0.002 | 0.003 |
| Science and technology studies | 0.001 | 0.004 |
| Scholarly communication | 0.003 | 0.006 |
| Open science | 0.006 | 0.001 |
| Research integrity | 0.017 | 0.016 |
| Insufficient payload (model declined to judge) | 0.011 | 0.004 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".