Comparing the responsiveness of a brief, multidimensional risk screening tool for back pain to its unidimensional reference standards: The whole is greater than the sum of its parts
Bibliographic record
Abstract
Back pain is a leading cause of disability. Previous research suggests that modifiable risk factors influence recovery from back pain, and practice guidelines recommend integrating such factors within primary care management. Toward this goal, a brief, multidimensional questionnaire, the STarT Back Tool, was designed to facilitate risk assessment by reducing the need to administer multiple, unidimensional questionnaires. However, aspects of this tool's clinical utility remain unaddressed. For instance, it is unclear whether this tool is responsive to treatment-related changes or whether clinically meaningful information is lost when it replaces multiple risk questionnaires. This study compared the responsiveness of the STarT Back Tool to its corresponding full-length measures, and evaluated its ability to detect clinically meaningful improvement. The study sample included 300 participants that consulted their doctor with disabling back pain. The STarT Back Tool and its reference standard questionnaires (disability, catastrophizing, fear, and depression) were administered at baseline and 4 months later. Regression analyses tested whether, after controlling for its reference standard questionnaires, the STarT Back Tool (independent variable) predicted treatment-related changes in global improvement, pain severity, disability, catastrophizing, fear, and depression (dependent variables). Receiver operating characteristic analyses determined the level of STarT Back change needed for clinically meaningful improvement. STarT Back scores predicted changes in all dependent variables except depression. Reductions in STarT Back scores predicted meaningful improvement on all dependent variables. These findings suggest that the STarT Back Tool, instead of multiple risk questionnaires, can be used to measure recovery from back pain. Implications for future research and clinical practice are discussed.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.022 | 0.006 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".