Derivation of clinical prediction rules for identifying patients with non-acute low back pain who respond best to a lumbar stabilization exercise program at post-treatment and six-month follow-up
Bibliographic record
Abstract
Low back pain (LBP) remains one of the most common and incapacitating health conditions worldwide. Clinical guidelines recommend exercise programs after the acute phase, but clinical effects are modest when assessed at a population level. Research needs to determine who is likely to benefit from specific exercise interventions, based on clinical presentation. This study aimed to derive clinical prediction rules (CPRs) for treatment success, using a lumbar stabilization exercise program (LSEP), at the end of treatment and at six-month follow-up. The eight-week LSEP, including clinical sessions and home exercises, was completed by 110 participants with non-acute LBP, with 100 retained at the six-month follow-up. Physical (lumbar segmental instability, motor control impairments, posture and range of motion, trunk muscle endurance and physical performance tests) and psychological (related to fear-avoidance and home-exercise adherence) measures were collected at a baseline clinical exam. Multivariate logistic regression models were used to predict clinical success, as defined by ≥50% decrease in the Oswestry Disability Index. CPRs were derived for success at program completion (T8) and six-month follow-up (T34), negotiating between predictive ability and clinical usability. The chosen CPRs contained four (T8) and three (T34) clinical tests, all theoretically related to spinal instability, making these CPRs specific to the treatment provided (LSEP). The chosen CPRs provided a positive likelihood ratio of 17.9 (T8) and 8.2 (T34), when two or more tests were positive. When applying these CPRs, the probability of treatment success rose from 49% to 96% at T8 and from 53% to 92% at T34. These results support the further development of these CPRs by proceeding to the validation stage.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".