The Ankylosing Spondylitis Performance Index: Reliability and Feasibility of an Objective Test for Physical Functioning
Bibliographic record
Abstract
OBJECTIVE: Physical function in patients with axial spondyloarthritis (axSpA) is currently evaluated through questionnaires. The Ankylosing Spondylitis Performance Index (ASPI) is a performance-based measure for physical functioning, which has been validated in Dutch patients with radiographic (r-) axSpA. The interrater reliability has not yet been determined. To our knowledge, this study is the first to evaluate the validity, reliability, and feasibility of the ASPI in another patient population, including both r- and nonradiographic (nr-) axSpA patients. METHODS: Patients with axSpA were recruited from rheumatology clinics in Santiago, Chile. Dutch instructions were translated to Spanish by a forward-backward procedure. Study visits were performed at baseline and 1-4 weeks later. Four ASPI observers were involved, measuring the performance times of the 3 ASPI tests. Validity was assessed through a patient questionnaire (numeric rating scale 0-10: ≥ 6 sufficient). For reliability, intraclass correlation coefficients (ICC) were calculated (with 95% CI). Correlations between the ASPI and disease variables were tested with regression analyses. RESULTS: Sixty-eight patients were included (57% male, 52% r-axSpA). All patients understood the Spanish instructions and considered the ASPI to reach its aim (84%) and representativeness (85%) for physical functioning. The overall interrater (n = 62) and test-retest (n = 39) reliability (ICC) of the 3 tests combined were 0.93 (0.88-0.96) and 0.94 (0.87-0.97), respectively. Eighty-two percent of the patients completed all tests and 94% finished in < 15 min (feasibility). CONCLUSION: This study demonstrated a high validity and feasibility in an entirely different population, with both r-axSpA and nr-axSpA. The interrater and test-retest reliability was excellent. The ASPI instructions are now available for Spanish-speaking patients.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.007 | 0.012 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.001 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".