The Ability of Scoliosis-Specific Patient-Reported Outcome Measures to Detect Change in Adolescents with Idiopathic Scoliosis
Bibliographic record
Abstract
Background: Adolescent Idiopathic Scoliosis (AIS) is a spinal disorder characterized by a lateral curvature greater than 10, with vertebral rotation. AIS is the most common form of scoliosis, affecting 1-3% of adolescents. Current AIS treatments are classified as either conservative (observation, bracing, exercise) or surgery. To assess the effect of treatments on health outcomes, clinicians often administer patient-reported outcome measures (PROMs). However, many scoliosis-specific PROMs, including the widely used Scoliosis Research Society-22r (SRS-22r), were developed for surgical populations and tend to exhibit ceiling effects when administered to patients with milder curves, limiting their ability to accurately reflect impacts on various health outcomes and their ability to detect change.Newer PROMs such as the Italian Spine Youth Quality of Life Questionnaire (ISYQOL), Body Image Disturbance Questionnaire–Scoliosis version (BIDQ-S), and Truncal Anterior Asymmetry Scoliosis Questionnaire (TAASQ) address limitations of traditional tools by reducing ceiling effects and incorporating constructs such as body image and anterior trunk appearance. While these PROMS show adequate reliability and validity, their responsiveness (i.e. ability to detect an important change), sensitivity to change (i.e. ability to detect change over and above measurement error), and patient-acceptable symptom state (PASS) (i.e. the value beyond which patients consider their health state as acceptable) are unknown.Objective: This study aimed to determine the responsiveness, sensitivity to change, and PASS of the ISYQOL, BIDQ-S, TAASQ, and SRS-22r in a sample of conservatively treated adolescents with idiopathic scoliosis.Methods: Adolescents with AIS undergoing conservative management were consecutively recruited from 13 physiotherapy and hospital-based scoliosis clinics across Canada and the United States. Participants completed the ISYQOL, BIDQ-S, TAASQ, SRS-22r, and a global rating of whether they perceived their condition to be acceptable (symptom-specific well-being item of the Core Outcome Measures Index (COMI)) at the initial visit and again at three and six months via REDCap. At three and six months, patients also completed a global rating of change (GRC). At physiotherapy clinics, participants received a Schroth exercise program, whereas, at hospital-based clinics, patients were either prescribed observation, referred for scoliosis-specific exercise, or prescribed bracing.The PASS thresholds and minimum clinically important difference (MCID) change scores were determined by first assessing the correlation between each of the PROM scores and the COMI question at baseline and the GRC scores (r >0.30) at three and six months. We then constructed receiver operating characteristic (ROC) curves for each PROM to identify the thresholds with the optimal ability to identify participants perceiving their condition as acceptable and as having experienced an important change. The area under the curve (AUC) was then measured, accepting the estimated PASS and MCID if AUC >0.70. Sensitivity to change for each PROM was measured by calculating standardized response means (SRM) with 95% confidence intervals.Results: The study included 28 participants (22 females, 6 males) with a mean age and curve angle of 13.5 ± 1.7 years and 32.6° ± 10.4°, respectively. PASS threshold estimates were identified for the ISYQOL (>47.9), ISYQOL International (>46), BIDQ-S (<1.9), several TAASQ domains/subdomains (<1.9 to <2.1), and most SRS-22r domains (>2.3 and >4.6). Six-month MCID change scores were identified for the ISYQOL (>-3.1), ISYQOL International (>-2.9), and several TAASQ domains/sub-domains (<0.0 to <0.8). At six months, small responsiveness was observed in the BIDQ-S (-0.39) and the TAASQ Appearance (-0.48), Clothing (-0.25), Clothing-General (-0.46), and Breast-Size (0.36) domains. Moderate to large responsiveness was recorded for the ISYQOL (0.68) and the ISYQOL International (1.17), respectively. Most PROMs showed improved sensitivity to change at six months compared to three, with the exception of the TAASQ Breast domain and sub-domains, which declined in performance.Conclusion: Once confirmed in a larger sample, the recommended PASS threshold and MCID change score estimates may provide clinicians with valuable insight into patient perspectives on treatment success. While definitive comparisons between the newer tools and the SRS-22r were limited by sample size, the ISYQOL International demonstrated promising responsiveness and sensitivity to change. The BIDQ-S and TAASQ were comparable, except for the TAASQ Breast domain, which showed a reduced ability to detect meaningful change. These findings support the further investigation of these PROMs, particularly the ISYQOL International, in the conservative treatment of AIS.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.014 | 0.040 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".