Patients and clinicians define symptom levels and meaningful change for PROMIS pain interference and fatigue in RA using bookmarking
Bibliographic record
Abstract
OBJECTIVES: Using patient-reported outcomes to inform clinical decision-making depends on knowing how to interpret scores. Patient-Reported Outcome Measurement Information System® (PROMIS®) instruments are increasingly used in rheumatology research and care, but there is little information available to guide interpretation of scores. We sought to identify thresholds and meaningful change for PROMIS Pain Interference and Fatigue scores from the perspective of RA patients and clinicians. METHODS: We developed patient vignettes using the PROMIS item banks representing a continuum of Pain Interference and Fatigue levels. During a series of face-to-face 'bookmarking' sessions, patients and clinicians identified thresholds for mild, moderate and severe levels of symptoms and identified change deemed meaningful for making treatment decisions. RESULTS: In general, patients selected higher cut points to demarcate thresholds than clinicians. Patients and clinicians generally identified changes of 5-10 points as representing meaningful change. The thresholds and meaningful change scores of patients were grounded in their lived experiences having RA, approach to self-management, and the impacts on function, roles and social participation. CONCLUSION: Results offer new information about how both patients and clinicians view RA symptoms and functional impacts. Results suggest that patients and providers may use different strategies to define and interpret RA symptoms, and select different thresholds when describing symptoms as mild, moderate or severe. The magnitude of symptom change selected by patients and clinicians as being clinically meaningful in interpreting treatment efficacy and loss of response may be greater than levels determined by external anchor and statistical methods.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".