American College of Rheumatology White Paper on Performance Outcome Measures in Rheumatology
Bibliographic record
Abstract
OBJECTIVE: To highlight the opportunities and challenges of developing and implementing performance outcome measures in rheumatology for accountability purposes. METHODS: We constructed a hypothetical performance outcome measure to demonstrate the benefits and challenges of designing quality measures that assess patient outcomes. We defined the data source, measure cohort, reporting period, period at risk, measure outcome, outcome attribution, risk adjustment, reliability and validity, and reporting approach. We discussed outcome measure challenges specific to rheumatology and to fields where patients have predominantly chronic, complex, ambulatory care-sensitive conditions. RESULTS: Our hypothetical outcome measure was a measure of rheumatoid arthritis disease activity intended for evaluating Accountable Care Organization performance. We summarized the components, benefits, challenges, and tradeoffs between feasibility and usability. We highlighted how different measure applications, such as for rapid cycle quality improvement efforts versus pay for performance programs, require different approaches to measure development and testing. We provided a summary table of key take-home points for clinicians and policymakers. CONCLUSION: Performance outcome measures are coming to rheumatology, and the most effective and meaningful measures can only be created through the close collaboration of patients, providers, measure developers, and policymakers. This study provides an overview of key issues and is intended to stimulate a productive dialogue between patients, practitioners, insurers, and government agencies regarding optimal performance outcome measure development.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.097 | 0.185 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.002 |
| Bibliometrics | 0.004 | 0.007 |
| Science and technology studies | 0.004 | 0.006 |
| Scholarly communication | 0.010 | 0.006 |
| Open science | 0.005 | 0.005 |
| Research integrity | 0.017 | 0.023 |
| Insufficient payload (model declined to judge) | 0.009 | 0.005 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".