Identifying Minimal and Meaningful Change in a <scp>Patient‐Reported</scp> Outcomes Measurement Information System for Rheumatoid Arthritis: Use of Multiple Methods and Perspectives
Bibliographic record
Abstract
Objective Rheumatoid arthritis (RA) is chronic, painful, disabling condition resulting in significant impairments in physical, emotional, and social health. Our objective was to use different methods and perspectives to evaluate the responsiveness of Patient‐Reported Outcomes Measurement Information System (PROMIS) short forms (SFs) and to identify minimal and meaningful score changes. Methods Adults with RA who were enrolled in a multisite prospective observational cohort completed PROMIS physical function, pain interference, fatigue, and participation in social roles/activities SFs, the PROMIS 29‐item form (PROMIS‐29), and pain and patient global assessment, and rated change in specific symptoms and RA (a little versus lot better or worse) at the second visit. Physicians recorded joint counts, physician global assessment, and change in RA at visit 2. We compared mean score differences for minimal and meaningful improvement/worsening using patient and physician change ratings and distribution‐based methods, and we visually inspected empirical cumulative distribution function curves by change categories. Results The 348 adults were mostly female (81%) with longstanding RA. Using patient ratings, generally 1–3‐point differences were observed for minimal change and 3–7 points for meaningful change. Larger differences were observed with patient versus physician ratings and for symptom‐specific versus RA change. Mean differences were similar among SF versions. Prespecified hypotheses about change in PROMIS physical function, pain interference, fatigue, and participation and legacy scales were supported. Conclusion PROMIS SFs and the PROMIS‐29 profiles are responsive to change and generally distinguish between minimal and meaningful improvement and worsening in key RA domains. These data add to a growing body of evidence demonstrating the robust psychometric properties of PROMIS and supporting its use in RA care, research, and decision‐making.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.008 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".