PERFORMANCE OF THE 2019 EULAR/ACR AND 2012 SLICC CLASSIFICATION CRITERIA FOR SLE AS DIAGNOSTIC CRITERIA
Bibliographic record
Abstract
PV240 / #328 Poster Topic: AS23 - SLE-Diagnosis, Manifestations, & Outcomes Background/Purpose 2019 EULAR/ACR SLE criteria were validated for classification from cases of established SLE compared to other diseases (sensitivity 0.96 (0.95 - 0.98); specificity 0.93 (0.91 - 0.95)),[1] but have not been assessed as diagnostic criteria in an undiagnosed cohort of ANA positive patients clinically suspected to have SLE. The aim is to evaluate, in these patients, the performance of 2019 EULAR/ACR SLE criteria against consultants’ diagnosis of SLE after 3 years follow-up and consultants’ treatment decision. Methods We included all consecutive consenting patients with ANA positive ≥ 1:80, new symptoms of ≤ 1 year, and SLE treatment naïve who had been referred to a specialist lupus clinic from primary care. 2019 EULAR/ACR SLE criteria and 2012 SLICC criteria were evaluated by research fellows at the time of the enrollment (T0) and evaluated annually, or on additional visits, during a 3-years follow-up. Clinical charts (at T0, and at the visit in which patients met classification criteria or last follow-up visit for those who did not meet them, T1) were anonymized and any documented diagnosis was redacted. Charts were reviewed by 4 consultants who had been independent of the patients’ clinical care to assess the diagnosis as well as confidence in diagnosis and suggested treatment decision. Interrater agreement was analyzed by Cohen’s k. Results 141 patients were included. Table 1 reports baseline characteristics. At T1, SLE was diagnosed by consultants in 26 patients (18.4%); a moderate interrater agreement was observed (k=0.52; CI 0.31–0.72). The 2019 EULAR/ACR classification criteria were met in 36 patients (26%) and the 2012 SLICC criteria in 33 patients (23.4%). Other consultants’ diagnoses were: Undifferentiated CTD (UCTD, 19.1%), Sjögren’s Disease (SjD, 8%), inflammatory arthritis (IA, 3.5%). Table 2 reports sensitivity, specificity, PPV, NPV of the 2019 EULAR/ACR SLE and 2012 SLICC criteria. At T1, in 21 patients (14.8%) SLE diagnosis was agreed by the consultants and 2019 EULAR/ACR SLE criteria; consultants’ suggested treatments for these patients were: immunosuppressant (IS) +/- hydroxychloroquine (HCQ) in 11 cases (52.4%), HCQ alone in 9 (42.8%), and no treatment in 1 (4.8%). In 5 patients (3.5%) SLE was diagnosed by consultants but not classified by the EULAR/ACR SLE criteria; the suggested treatment from the consultants were: IS in 3 cases (60%), HCQ in 2 (40%). In 15 patients where 2019 EULAR/ACR criteria were met, consultants did not diagnose SLE (10.6%); consultants’ diagnoses were: SjD (40%), UCTD (33.3%), IA (26.7%%); treatment were: 7 HCQ (46.7%; in 1 case with im steroid), 5 IS (33.3%), 1 im steroid (6.7%) and no treatment in 2 (13%). In addition to the analyses above, at T0, the consultants diagnosed 11 patients who did not meet EULAR/ACR criteria as SLE. But at T1, 6/11 (55%) of these were subsequently classified as SLE. Thus, SLE could be diagnosed earlier than using classification criteria in some patients. Table 1: Features at baseline Table 2: Sensitivity, Specificity, Positive Predictive Value, Negative Predictive Value Conclusions When used in a diagnostic setting, the performance of 2019 EULAR/ACR SLE criteria was less good than for classification. Patients meeting classification criteria were usually clinically diagnosed as SLE as well, but some patients with a consultant diagnosis of SLE did not meet criteria and in these cases, treatment decisions were similar to those meeting criteria. Classification criteria should be used with caution for the diagnosis when evaluating patients with suspected SLE. Future work will analyze a second cohort. References: [1.] Aringer M. ARD 2019;78(9):1151-9.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".