VALIDATION OF AMERICAN COLLEGE OF RHEUMATOLOGY DIGITAL CLINICAL QUALITY MEASURES FOR LUPUS CARE IN THE RHEUMATOLOGY INFORMATICS SYSTEM FOR EFFECTIVENESS (RISE) REGISTRY
Bibliographic record
Abstract
PV086 / #369 Poster Topic: AS11 - Epidemiology and Public Health Background/Purpose In collaboration with the Centers for Disease Control and Prevention (CDC), the American College of Rheumatology (ACR) recently developed digital clinical quality measures for lupus clinical care. These measures included 1) hydroxychloroquine use (yes/no), 2) limiting glucocorticoid use of doses above 7.5 mg/day of prednisone to fewer than 6 months (yes/no), and 3) kidney function and urine protein laboratory testing for monitoring/screening for lupus nephritis (yes/no). We aimed to assess the accuracy of calculating these quality measures in the ACR’s Rheumatology Informatics System for Effectiveness (RISE) registry compared to manual medical record review. Methods Five practices that participate in the RISE registry were recruited to participate in this study. All eligible patients with systemic lupus erythematosus (SLE) (defined as ≥ 2 SLE codes ≥ 30 days apart in 2022) who were at least 18 years of age and had at least 1 visit with a participating practice in 2022 were identified from the RISE registry. Among those patients, stratified random sampling was used to select 35 patients from each practice based on patients’ glucocorticoid and hydroxychloroquine use in 2022. Practice staff were asked to complete a structured medical record review for at least 25 of the 35 patients at each site. Data related to the SLE quality measures were abstracted for the measurement year 2022, including hydroxychloroquine use and its contraindications, glucocorticoid use, including dose and duration of use, end-stage kidney disease (ESKD), and laboratory testing for kidney function and urine protein excretion. Corresponding electronic health record data on these patients were extracted from the RISE registry. Cohen’s Kappa statistic and percent agreement were calculated for each measure component. Results We included 121 patients with SLE, of whom 92% were female, 21% were Black or African American, and 13% had lupus nephritis (Table 1). Agreement between the practice medical record review and RISE data varied across the measure components, with higher Kappa statistics and percent agreement for medication use than for laboratory testing (Table 2). The Kappa for hydroxychloroquine use and glucocorticoid use were 0.70 (95% CI 0.60- 0.90) and 0.90 (95% CI 0.82-0.98), respectively. Glucocorticoid use over 7.5 mg/day for longer than 6 months was infrequent, with 97% agreement, but lower kappa 0.48 (95% CI (0.05-0.92). One patient was identified with ESKD, with 100% agreement between data sources, and was excluded from the kidney monitoring measure. Kidney function measurement was reported at least once in 2022 for 93% per medical record review and 83% of patients per RISE data, with Kappa 0.27 (95% CI 0.04-0.50). Urine protein measurement was reported at least once in 2022 for 72% per medical record review and 49% of patients per RISE data, with Kappa 0.35 (95% CI 0.21-0.50). Table 1. Characteristics of Patients with Systemic Lupus Erythematosus Table 2. Agreement between RISE data and Medical Record Review Conclusions In this initial validation study of digital clinical quality measures for lupus applied to 5 rheumatology practices in the RISE registry, we found good agreement for hydroxychloroquine and glucocorticoid use between medical record review and RISE data. Discrepancies in glucocorticoid dosing rarely led to misclassification for the glucocorticoid measure. Agreement was lower for the capture of kidney monitoring tests. Further work will include additional measure validity testing between RISE and Medicare data and will examine strategies to more accurately capture kidney monitoring for patients with lupus.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.136 | 0.249 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.004 | 0.005 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.003 | 0.002 |
| Open science | 0.002 | 0.003 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".