Derivation of clinical prediction rules for identifying patients with non-acute low back pain who respond best to a lumbar stabilization exercise program at post-treatment and six-month follow-up
Bibliographic record
Abstract
Low back pain (LBP) remains one of the most common and incapacitating health conditions worldwide. Clinical guidelines recommend exercise programs after the acute phase, but clinical effects are modest when assessed at a population level. Research needs to determine who is likely to benefit from specific exercise interventions, based on clinical presentation. This study aimed to derive clinical prediction rules (CPRs) for treatment success, using a lumbar stabilization exercise program (LSEP), at the end of treatment and at six-month follow-up. The eight-week LSEP, including clinical sessions and home exercises, was completed by 110 participants with non-acute LBP, with 100 retained at the six-month follow-up. Physical (lumbar segmental instability, motor control impairments, posture and range of motion, trunk muscle endurance and physical performance tests) and psychological (related to fear-avoidance and home-exercise adherence) measures were collected at a baseline clinical exam. Multivariate logistic regression models were used to predict clinical success, as defined by ≥50% decrease in the Oswestry Disability Index. CPRs were derived for success at program completion (T8) and six-month follow-up (T34), negotiating between predictive ability and clinical usability. The chosen CPRs contained four (T8) and three (T34) clinical tests, all theoretically related to spinal instability, making these CPRs specific to the treatment provided (LSEP). The chosen CPRs provided a positive likelihood ratio of 17.9 (T8) and 8.2 (T34), when two or more tests were positive. When applying these CPRs, the probability of treatment success rose from 49% to 96% at T8 and from 53% to 92% at T34. These results support the further development of these CPRs by proceeding to the validation stage.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.006 | 0.035 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".