MétaCan
Menu
Back to cohort

PERFORMANCE OF THE 2019 EULAR/ACR AND 2012 SLICC CLASSIFICATION CRITERIA FOR SLE AS DIAGNOSTIC CRITERIA

2025· article· en· W4410513050 on OpenAlexvenueno aff
Francesca Crisafulli, Md Yuzaiful Md Yusof, Aamir Aslam, Andrew Barr, Lesley-Anne Bissell, Shouvik Dass, Jack Arnold, Porntip Intapiboon, Franco Franceschini, Edward M Vital

Bibliographic record

VenueThe Journal of Rheumatology · 2025
Typearticle
Languageen
FieldEnvironmental Science
TopicAir Quality Monitoring and Forecasting
Canadian institutionsnot available
Fundersnot available
KeywordsMedicinePhysical therapyInternal medicine

Abstract

fetched live from OpenAlex

PV240 / #328 Poster Topic: AS23 - SLE-Diagnosis, Manifestations, & Outcomes Background/Purpose 2019 EULAR/ACR SLE criteria were validated for classification from cases of established SLE compared to other diseases (sensitivity 0.96 (0.95 - 0.98); specificity 0.93 (0.91 - 0.95)),[1] but have not been assessed as diagnostic criteria in an undiagnosed cohort of ANA positive patients clinically suspected to have SLE. The aim is to evaluate, in these patients, the performance of 2019 EULAR/ACR SLE criteria against consultants’ diagnosis of SLE after 3 years follow-up and consultants’ treatment decision. Methods We included all consecutive consenting patients with ANA positive ≥ 1:80, new symptoms of ≤ 1 year, and SLE treatment naïve who had been referred to a specialist lupus clinic from primary care. 2019 EULAR/ACR SLE criteria and 2012 SLICC criteria were evaluated by research fellows at the time of the enrollment (T0) and evaluated annually, or on additional visits, during a 3-years follow-up. Clinical charts (at T0, and at the visit in which patients met classification criteria or last follow-up visit for those who did not meet them, T1) were anonymized and any documented diagnosis was redacted. Charts were reviewed by 4 consultants who had been independent of the patients’ clinical care to assess the diagnosis as well as confidence in diagnosis and suggested treatment decision. Interrater agreement was analyzed by Cohen’s k. Results 141 patients were included. Table 1 reports baseline characteristics. At T1, SLE was diagnosed by consultants in 26 patients (18.4%); a moderate interrater agreement was observed (k=0.52; CI 0.31–0.72). The 2019 EULAR/ACR classification criteria were met in 36 patients (26%) and the 2012 SLICC criteria in 33 patients (23.4%). Other consultants’ diagnoses were: Undifferentiated CTD (UCTD, 19.1%), Sjögren’s Disease (SjD, 8%), inflammatory arthritis (IA, 3.5%). Table 2 reports sensitivity, specificity, PPV, NPV of the 2019 EULAR/ACR SLE and 2012 SLICC criteria. At T1, in 21 patients (14.8%) SLE diagnosis was agreed by the consultants and 2019 EULAR/ACR SLE criteria; consultants’ suggested treatments for these patients were: immunosuppressant (IS) +/- hydroxychloroquine (HCQ) in 11 cases (52.4%), HCQ alone in 9 (42.8%), and no treatment in 1 (4.8%). In 5 patients (3.5%) SLE was diagnosed by consultants but not classified by the EULAR/ACR SLE criteria; the suggested treatment from the consultants were: IS in 3 cases (60%), HCQ in 2 (40%). In 15 patients where 2019 EULAR/ACR criteria were met, consultants did not diagnose SLE (10.6%); consultants’ diagnoses were: SjD (40%), UCTD (33.3%), IA (26.7%%); treatment were: 7 HCQ (46.7%; in 1 case with im steroid), 5 IS (33.3%), 1 im steroid (6.7%) and no treatment in 2 (13%). In addition to the analyses above, at T0, the consultants diagnosed 11 patients who did not meet EULAR/ACR criteria as SLE. But at T1, 6/11 (55%) of these were subsequently classified as SLE. Thus, SLE could be diagnosed earlier than using classification criteria in some patients. Table 1: Features at baseline Table 2: Sensitivity, Specificity, Positive Predictive Value, Negative Predictive Value Conclusions When used in a diagnostic setting, the performance of 2019 EULAR/ACR SLE criteria was less good than for classification. Patients meeting classification criteria were usually clinically diagnosed as SLE as well, but some patients with a consultant diagnosis of SLE did not meet criteria and in these cases, treatment decisions were similar to those meeting criteria. Classification criteria should be used with caution for the diagnosis when evaluating patients with suspected SLE. Future work will analyze a second cohort. References: [1.] Aringer M. ARD 2019;78(9):1151-9.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.001
metaresearch head score (Gemma)0.001
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: Observational
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.042
Threshold uncertainty score0.163

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0010.001
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0000.000
Bibliometrics0.0000.000
Science and technology studies0.0000.000
Scholarly communication0.0000.000
Open science0.0000.000
Research integrity0.0000.000
Insufficient payload (model declined to judge)0.0000.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.023
GPT teacher head0.299
Teacher spread0.276 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designObservational
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2025
Admission routes1
Has abstractyes

Explore more

Same venueThe Journal of RheumatologySame topicAir Quality Monitoring and ForecastingFrench-language works237,207