Comparison of the 2019 European Alliance of Associations for Rheumatology/American College of Rheumatology Systemic Lupus Erythematosus Classification Criteria With Two Sets of Earlier Systemic Lupus Erythematosus Classification Criteria
Bibliographic record
Abstract
OBJECTIVE: The Systemic Lupus International Collaborating Clinics (SLICC) 2012 systemic lupus erythematosus (SLE) classification criteria and the revised American College of Rheumatology (ACR) 1997 criteria are list based, counting each SLE manifestation equally. We derived a classification rule based on giving variable weights to the SLICC criteria and compared its performance to the revised ACR 1997, the unweighted SLICC 2012, and the newly reported European Alliance of Associations for Rheumatology (EULAR)/ACR 2019 criteria sets. METHODS: The physician-rated patient scenarios used to develop the SLICC 2012 classification criteria were reemployed to devise a new weighted classification rule using multiple linear regression. The performance of the rule was evaluated on an independent set of expert-diagnosed patient scenarios and compared to the performance of the previously reported classification rules. RESULTS: The weighted SLICC criteria and the EULAR/ACR 2019 criteria had less sensitivity but better specificity compared to the list-based revised ACR 1997 and SLICC 2012 classification criteria. There were no statistically significant differences between any pair of rules with respect to overall agreement with the physician diagnosis. CONCLUSION: The 2 new weighted classification rules did not perform better than the existing list-based rules in terms of overall agreement on a data set originally generated to assess the SLICC criteria. Given the added complexity of summing weights, researchers may prefer the unweighted SLICC criteria. However, the performance of a classification rule will always depend on the populations from which the cases and non-cases are derived and whether the goal is to prioritize sensitivity or specificity.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.016 | 0.061 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.005 | 0.002 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".