Methodology for the Systematic Reviews on an Evidence-Based Approach for the Management of Chronic Low Back Pain
Bibliographic record
Abstract
STUDY DESIGN: Systematic review. OBJECTIVE: To provide a detailed description of the methods undertaken in the systematic search and analytical summary of chronic low back pain (CLBP) management issues and to describe the process used to develop clinical recommendations regarding challenges in the management of patients with CLBP. SUMMARY OF BACKGROUND DATA: We present methods used in conducting the systematic, evidence-based reviews and development of expert panel recommendations on key challenges to CLBP assessment and management. Our intent is that clinicians will combine the information from these reviews with an understanding of their own capacities and experience to better manage patients with chronic LBP and to consider future research that identifies patients or subgroups that respond differently with regard to benefits and safety to various treatment interventions. METHODS: A systematic search and critical review of the English language literature was undertaken for articles published on the classification, measurement, and management of CLBP. Citations were screened for relevance using a priori criteria, and relevant studies were critically reviewed. Whether an article was included for review depended on whether the study question was descriptive, one of therapy, one of prognosis, or one of diagnosis. When evaluating differential treatment benefits by specific disease, sociodemographic, and psychological subgroups, we sought to evaluate the heterogeneity of treatment effects. Studies were included if they made the treatment comparison and presented treatment effects by the predefined subgroup. The strength of evidence for the overall body of literature in each topic area was determined by two independent reviewers considering risk of bias, consistency, directness, and precision of results using a modification of the Grades of Recommendation Assessment, Development and Evaluation (GRADE) criteria. Disagreements were resolved by consensus. Findings from studies meeting inclusion criteria were summarized. From these summaries, clinical recommendations were formulated from consensus achieved among subject experts through a modified Delphi process. RESULTS: We identified and screened 2845 citations in 13 topic areas relating to the classification, measurement, and management of CLBP. Of these, 118 met our predetermined inclusion criteria and were used to attempt to answer specific clinical questions within each topic area. Some of the highlights of the analysis revealed a limited number of studies meeting inclusion criteria for topics evaluating therapy, use of magnetic resonance imaging, and classification systems. Few studies comparing surgical fusion to nonoperative care were identified that presented treatment effects by subgroups limiting the evaluation of heterogeneity of treatment effects. CONCLUSION: We undertook systematic reviews to understand the classification, measurement, and management of CLBP and to provide clinical recommendations. This article reports the methods used in the reviews. CLINICAL RECOMMENDATIONS: Clinical recommendations were made where appropriate using the GRADE/Agency for Healthcare Research and Quality approach, which imparts a deliberate separation between the quality of the evidence (i.e., high, moderate, low, or inconclusive) from the strength of the recommendation. The quality of evidence plays only a part as the strength of the recommendation reflects the extent to which we can, across the range of patients for whom the recommendations are intended, be confident that desirable effects of a management strategy outweigh undesirable effects.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.333 | 0.476 |
| Meta-epidemiology (narrow) | 0.007 | 0.006 |
| Meta-epidemiology (broad) | 0.024 | 0.032 |
| Bibliometrics | 0.045 | 0.035 |
| Science and technology studies | 0.004 | 0.006 |
| Scholarly communication | 0.013 | 0.010 |
| Open science | 0.010 | 0.011 |
| Research integrity | 0.011 | 0.010 |
| Insufficient payload (model declined to judge) | 0.037 | 0.009 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".