Physical functioning in the lumbar spinal surgery population: A systematic review and narrative synthesis of outcome measures and measurement properties of the physical measures
Bibliographic record
Abstract
BACKGROUND: International agreement supports physical functioning as a key domain to measure interventions effectiveness for low back pain. Patient reported outcome measures (PROMs) are commonly used in the lumbar spinal surgery population but physical functioning is multidimensional and necessitates evaluation also with physical measures. OBJECTIVE: 1) To identify outcome measures (PROMs and physical) used to evaluate physical functioning in the lumbar spinal surgery population. 2) To assess measurement properties and describe the feasibility and interpretability of physical measures of physical functioning in this population. STUDY DESIGN: Two-staged systematic review and narrative synthesis. METHODS: This systematic review was conducted according to a registered and published protocol. Two stages of searching were conducted in MEDLINE, EMBASE, Health & Psychosocial Instruments, CINAHL, Web of Science, PEDro and ProQuest Dissertations & Theses. Stage one included studies to identify physical functioning outcome measures (PROMs and physical) in the lumbar spinal surgery population. Stage two (inception to 10 July 2023) included studies assessing measurement properties of stage one physical measures. Two independent reviewers determined study eligibility, extracted data and assessed risk of bias (RoB) according to COSMIN guidelines. Measurement properties were rated according to COSMIN criteria. Level of evidence was determined using a modified GRADE approach. RESULTS: Stage one included 1,101 reports using PROMs (n = 70 established in literature, n = 67 developed by study authors) and physical measures (n = 134). Stage two included 43 articles assessing measurement properties of 34 physical measures. Moderate-level evidence supported sufficient responsiveness of 1-minute stair climb and 50-foot walk tests, insufficient responsiveness of 5-minute walk and sufficient reliability of distance walked during the 6-minute walk. Very low/low-level evidence limits further understanding. CONCLUSIONS: Many physical measures of physical functioning are used in lumbar spinal surgery populations. Few have investigations of measurement properties. Strongest evidence supports responsiveness of 1-minute stair climb and 50-foot walk tests and reliability of distance walked during the 6-minute walk. Further recommendations cannot be made because of very low/low-level evidence. Results highlight promise for a range of measures, but prospective, low RoB studies are required.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.046 | 0.156 |
| Meta-epidemiology (narrow) | 0.002 | 0.002 |
| Meta-epidemiology (broad) | 0.010 | 0.008 |
| Bibliometrics | 0.020 | 0.017 |
| Science and technology studies | 0.002 | 0.002 |
| Scholarly communication | 0.005 | 0.005 |
| Open science | 0.002 | 0.003 |
| Research integrity | 0.003 | 0.002 |
| Insufficient payload (model declined to judge) | 0.004 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".