Prioritizing Patient-Reported Outcome Measures for Routine Collection in Rheumatoid Arthritis: An Integrated Consensus-Building Process with Patients and Health Care Providers
Bibliographic record
Abstract
Objectives Patient-reported outcome measures (PROMs) are essential tools for prioritizing patient-centered care. Effective use of PROMs collected in clinical care hinges on selecting appropriate measures. The aim of this study was to establish consensus on PROMs for routine collection in RA care based on Canadian patient and provider preferences. Methods Candidate PROMs were identified through an environmental scan. Candidate PROMs met the following criteria: 1) clinical and/or research evidence supporting their use, 2) valid psychometric properties, 3) available at no cost, and 4) feasible to complete in routine care. We adopted a dual-panel consensus-building approach with a patient-exclusive panel, and separate engagement with rheumatology healthcare providers (HCPs). Consensus among patients with RA was established via a modified Delphi, co-led with 2 patient partners. This began with a prioritization of health domains/outcomes for review (Round 0). PROMs were then rated using a Likert scale of 1-9 for content validity, importance, and feasibility (Round 1), followed by a virtual consensus-building discussion (Round 2), and final re-rating of PROMs (Round 3). HCPs were engaged in a two-part rating and discussion exercise with 2 groups of providers from Alberta academic centers. Consensus was achieved if PROMs received a median rating of 7 or higher in at least 2 criteria. We included a ranking exercise for highly rated PROMs (>7 in all 3 criteria) in the final patient panel round. An overall rank was assigned to PROMs based on the rank-weighted sum. Results Fifteen patients with RA took part in the Delphi rounds. From an initial set of 15 candidate PROMS, 10 were rated highly (≥ 7) across all 3 criteria in Round 3. These PROMs were subsequently ranked. A measure of physical function ranked highest. Among HCPs (n=23), 4 PROMs received a median rating of 7 or higher in at least 2 criteria. Though limited, there was agreement between patients and HCPs ratings for pain and fatigue PROMs. Of these, only pain interference was highly ranked by patients (Table 1). Feedback from providers suggests that PROMs should align with clinical documentation needs, including insurance forms for advanced therapies and disability claims. Table 1: Patients and Health Care Providers (HCPs) PROM consensus results Conclusion Given the limited agreement between patients and providers on PROM measurement priorities, further work is required to understand why and to select a meaningful set of PROMs for routine collection. To start, a measure of physical function will be initialized for routine collection as it is a high patient priority and supports HCPs clinical documentation.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.418 | 0.320 |
| Meta-epidemiology (narrow) | 0.002 | 0.002 |
| Meta-epidemiology (broad) | 0.003 | 0.003 |
| Bibliometrics | 0.012 | 0.007 |
| Science and technology studies | 0.007 | 0.003 |
| Scholarly communication | 0.007 | 0.006 |
| Open science | 0.006 | 0.015 |
| Research integrity | 0.003 | 0.007 |
| Insufficient payload (model declined to judge) | 0.003 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".