Roland-Morris Disability Questionnaire and Oswestry Disability Index: Which Has Better Measurement Properties for Measuring Physical Functioning in Nonspecific Low Back Pain? Systematic Review and Meta-Analysis
Bibliographic record
Abstract
BACKGROUND: Physical functioning is a core outcome domain to be measured in nonspecific low back pain (NSLBP). A panel of experts recommended the Roland-Morris Disability Questionnaire (RMDQ) and Oswestry Disability Index (ODI) to measure this domain. The original 24-item RMDQ and ODI 2.1a are recommended by their developers. PURPOSE: The purpose of this study was to evaluate whether the 24-item RMDQ or the ODI 2.1a has better measurement properties than the other to measure physical functioning in adult patients with NSLBP. DATA SOURCES: Bibliographic databases (MEDLINE, Embase, CINAHL, SportDiscus, PsycINFO, and Google Scholar), references of existing reviews, and citation tracking were the data sources. STUDY SELECTION: Two reviewers selected studies performing a head-to-head comparison of measurement properties (reliability, validity, and responsiveness) of the 2 questionnaires. The COnsensus-based Standards for the selection of health Measurement INstruments (COSMIN) checklist was used to assess the methodological quality of these studies. DATA EXTRACTION: The studies' characteristics and results were extracted by 2 reviewers. A meta-analysis was conducted when there was sufficient clinical and methodological homogeneity among studies. DATA SYNTHESIS: Nine articles were included, for a total of 11 studies assessing 5 measurement properties. All studies were classified as having poor or fair methodological quality. The ODI displayed better test-retest reliability and smaller measurement error, whereas the RMDQ presented better construct validity as a measure of physical functioning. There was conflicting evidence for both instruments regarding responsiveness and inconclusive evidence for internal consistency. LIMITATIONS: The results of this review are not generalizable to all available versions of these questionnaires or to patients with specific causes for their LBP. CONCLUSIONS: Based on existing head-to-head comparison studies, there are no strong reasons to prefer 1 of these 2 instruments to measure physical functioning in patients with NSLBP, but studies of higher quality are needed to confirm this conclusion. Foremost, content, structural, and cross-cultural validity of these questionnaires in patients with NSLBP should be assessed and compared.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.048 | 0.121 |
| Meta-epidemiology (narrow) | 0.003 | 0.002 |
| Meta-epidemiology (broad) | 0.032 | 0.037 |
| Bibliometrics | 0.014 | 0.012 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.007 | 0.006 |
| Open science | 0.003 | 0.002 |
| Research integrity | 0.004 | 0.003 |
| Insufficient payload (model declined to judge) | 0.005 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".