Development and Initial Validation of the Back Pain Functional Scale
Bibliographic record
Abstract
STUDY DESIGN: A prospective repeated-measures design was applied. OBJECTIVES: To examine the measurement properties of the Back Pain Functional Scale (BPFS) and the Roland-Morris Questionnaire (RMQ) and to formulate hypotheses and sample size estimates for a subsequent comparison study. SUMMARY OF BACKGROUND DATA: Although there are numerous functional status measures for patients with low back pain, most have been conceived of and validated with a group rather than an individual patient as the unit of interest. Also, little has been done to formally compare-this includes the generation of a priori hypotheses, followed by statistical hypotheses testing-the many competing measures. METHODS: Subjects were 77 patients with low back pain who were referred by physicians to 10 outpatient physical therapy clinics located in Canada and the United States. The questionnaires were administered at patients' initial visits, within 48 hours of the initial visit, and at 1-, 2-, and 3-week follow-up visits. Reliability, cross-sectional validity, and longitudinal validity (sensitivity to change) coefficients were calculated. RESULTS: Test-retest reliability estimates of 0.81 and 0. 88 were obtained for the RMQ and BPFS, respectively. The measures demonstrated similar levels of cross-sectional validity. Correlations of 0.56 and 0.65 were noted between a prognostic rating of change and the RMQ and BPFS, respectively. The RMQ demonstrated a ceiling effect. Approximately 180 patients are needed for a subsequent head-to-head comparison study of the measures. CONCLUSIONS: The BPFS appears to have sound measurement properties, and a formal head-to-head comparison study with the RMQ is warranted.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".