The Performances of the ACR 1997, SLICC 2012, and EULAR/ACR 2019 Classification Criteria in Pediatric Systemic Lupus Erythematosus
Bibliographic record
Abstract
OBJECTIVE: Systemic lupus erythematosus (SLE) is a chronic autoimmune disease. The American College of Rheumatology (ACR) 1997, Systemic Lupus International Collaborating Clinics (SLICC) 2012, and European League Against Rheumatism (EULAR)/ACR 2019 SLE classification criteria are formed based on data mainly from adult patients. We aimed to test the performances of the SLE classification criteria among pediatric patients with SLE. METHODS: Pediatric patients with SLE (n = 262; 80.9% female) were included from 3 different centers in Turkey. As controls, 174 children (60.9% female) with other diseases who had ANA (antinuclear antibody) test results were included. The gold standard for SLE diagnosis was expert opinion. RESULTS: The sensitivities of the ACR 1997, SLICC 2012, and EULAR/ACR 2019 criteria were 68.7%, 95.4%, and 91.6%, respectively. The specificities of the ACR 1997, SLICC 2012, and EULAR/ACR 2019 criteria were 94.8%, 89.7%, and 88.5%, respectively. Eighteen patients with SLE met the SLICC 2012 but not the EULAR/ACR 2019 criteria. Among these, hematologic involvement was prominent (n = 13; 72.2%). Eight patients with SLE fulfilled the EULAR/ACR 2019 but not the SLICC 2012 criteria. Among these, joint involvement was prominent (n = 6; 75%). CONCLUSION: To our knowledge, this is the largest cohort study of pediatric SLE to test the performances of all 3 classification criteria. The SLICC 2012 criteria yielded the best sensitivity, whereas the ACR 1997 criteria had the best specificity. SLICC 2012 criteria performed better than EULAR/ACR 2019 criteria. Separation of different hematological manifestations in the SLICC 2012 criteria might have contributed to the higher performance of this criteria set.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".