Prioritizing patient-reported outcome measures for routine collection in rheumatoid arthritis: An integrated consensus-building process with patients and health care providers
Bibliographic record
Abstract
Background Effective clinical use of patient-reported outcome measures (PROMs) hinges on selecting appropriate measures. The aim of this study was to establish priorities for PROM routine collection in RA care at a health-system level based on patient and provider preferences. Methods Candidate PROMs were identified through an environmental scan. We adopted a dual-panel consensus-building approach with a patient-exclusive panel, and separate engagement with rheumatology healthcare providers (HCPs). Consensus among patients with RA was established via a modified Delphi process, co-led with two patient partners. HCPs were engaged in a rating and discussion exercise with two groups of providers from academic centres. Patient-panel consensus was achieved if median ratings on a 9-point Likert scale were ≥7 in all 3 criteria (importance, content validity and feasibility) and results informed HCP discussions for final selection. Results 15 patients with RA participated in the Delphi rounds and 22 HCPs participated in a rating exercise followed by two group discussions attended by a total of 43 rheumatologists. From an initial set of 15 candidate PROMs, 9 were rated highly (≥ 7) across all three criteria in Round 3 by patients. Among these, the PROMIS Physical Function 10a ranked highest. HCP feedback suggests a preference for PROMs that align with clinical documentation needs, including insurance forms for advanced therapies and disability claims. Conclusion A measure of physical function achieved a high priority from the patient panel, and it also supports clinical documentation needs of HCPs, and will thus be an initial implementation focus.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.452 | 0.284 |
| Meta-epidemiology (narrow) | 0.002 | 0.002 |
| Meta-epidemiology (broad) | 0.003 | 0.003 |
| Bibliometrics | 0.009 | 0.004 |
| Science and technology studies | 0.006 | 0.004 |
| Scholarly communication | 0.007 | 0.008 |
| Open science | 0.005 | 0.016 |
| Research integrity | 0.004 | 0.007 |
| Insufficient payload (model declined to judge) | 0.003 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".