Evaluating Common Outcomes for Measuring Treatment Success for Chronic Low Back Pain
Bibliographic record
Abstract
STUDY DESIGN: Systematic review. OBJECTIVE: To identify, describe, and evaluate common outcome measures in patients with chronic low back pain (CLBP). SUMMARY OF BACKGROUND DATA: The treatment of CLBP has been associated with multiple clinical challenges. Further complicating this is the myriad of outcome scores used to assess treatment of CLBP. These scores have been used to examine different domains of patient satisfaction and quality of life in the literature. Critical assessment of the frequency, parity, and the quality of these outcomes are essential to improve our understanding of CLBP. METHODS: A systematic review of the English-language literature was undertaken for articles published from January 2001 through December 31, 2010. Electronic databases and reference lists of key articles were searched to identify measures used to evaluate outcomes in six different domains in patients with CLBP. The titles and abstracts of the peer-reviewed literature of LBP were searched to determine which of these measures were most commonly reported in the literature and which have been validated in populations with CLBP. RESULTS: We identified 75 outcome measures cited to evaluate CLBP. Twenty-nine of these outcome measures were excluded because of only a single citation leaving 46 measures for the evaluation. The most commonly used functional outcomes were the Oswestry Disability Index, Roland Morris Disability Index, and range of motion. For pain, the Numeric Pain Rating Scale, Brief Pain Inventory, Pain Disability Index, McGill Pain Questionnaire, and visual analog scale were most commonly cited. For psychosocial function, the Fear Avoidance Beliefs Questionnaire, Tampa Scale for Kinesiophobia, and Beck Depression Inventory were most commonly used. For generic quality of life, short form 36, Nottingham Health Profile, short form 12, and Sickness Impact Profile were the most common measures. For objective measures, the work status/return to work, complications or adverse events, and medications used were the most commonly cited. For preference-based measures, the Euro-Quol 5 dimensions and short form 6 dimensions were most commonly cited. The validity, reliability, responsiveness, universality, and potential proprietary requirements are summarized for each. CONCLUSION: Outcome measures should be routinely assessed in patients with CLBP. The choice of appropriate outcome measure should be influenced by the study objectives and design, as well as properties of the particular measure within the context of CLBP. CLINICAL RECOMMENDATIONS: Recommendation 1: When selecting the appropriate outcome measures for clinical or research purposes, consider domains that best measure what are most important to patients. Measures that are valid, reliable, and responsive to change should be considered first. Other considerations include the number of items required (especially in the context of multiple measures), whether the measure is validated in the relevant language, and the associated costs or fees. Strength: Strong Recommendation 2: Domains of greatest importance include pain, function, and quality of life. If cost utilization is a priority, then preference-based measures should be considered. For pain, we recommend the VAS and NRPS because of their ease of administration and responsiveness. For function, we recommend the ODI and RMDQ. The SF-36 and its shorter versions are most commonly used and should be considered if quality of life is important. If cost utility is important, consider the EQ-5D or SF-6D. Psychosocial tests are best used as screening tools prior to surgery because of their lack of responsiveness. Complications should always be assessed as a standard of clinical practice. Return to work and medication use are complicated outcome measures and not recommended unless the specific study question is focused on these domains. Consider staff and patient burden when prioritizing one's battery of measures.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.053 | 0.175 |
| Meta-epidemiology (narrow) | 0.003 | 0.001 |
| Meta-epidemiology (broad) | 0.016 | 0.016 |
| Bibliometrics | 0.018 | 0.018 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.005 | 0.004 |
| Open science | 0.003 | 0.003 |
| Research integrity | 0.003 | 0.002 |
| Insufficient payload (model declined to judge) | 0.004 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".