Utility of the American-European Consensus Group and American College of Rheumatology Classification Criteria for Sjogren's syndrome in patients with systemic autoimmune diseases in the clinical setting
Bibliographic record
Abstract
OBJECTIVE: The aim of this study was to evaluate the feasibility and performance of the American-European Consensus Group (AECG) and ACR Classification Criteria for SS in patients with systemic autoimmune diseases. METHODS: Three hundred and fifty patients with primary SS, SLE, RA or scleroderma were randomly selected from our patient registry. Each patient was clinically diagnosed as probable/definitive SS or non-SS following a standardized evaluation including clinical symptoms and manifestations, confirmatory tests, fluorescein staining test, autoantibodies, lip biopsy and medical chart review. Using the clinical diagnosis as the gold standard, the degree of agreement with each criteria set and between the criteria sets was estimated. RESULTS: One hundred fifty-four (44%) patients were diagnosed with SS. The AECG criteria were incomplete in 36 patients (10.3%) and the ACR criteria in 96 (27.4%; P < 0.001). Nevertheless, their ability to classify patients was almost identical, with a sensitivity of 61.6 vs 62.3 and a specificity of 94.3 vs 91.3, respectively. Either set of criteria was met by 123 patients (80%); 95 (61.7%) met the AECG criteria and 96 (62.3%) met the ACR criteria, but only 68 (44.2%) patients met both sets. The concordance rate between clinical diagnosis and AECG or ACR criteria was moderate (k statistic 0.58 and 0.55, respectively). Among 99 patients with definitive SS sensitivity was 83.3 vs 77.7 and specificity was 90.8 vs 85.6, respectively. A discrepancy between clinical diagnosis and criteria was seen in 59 patients (17%). CONCLUSION: The feasibility of the SS AECG criteria is superior to that of the ACR criteria, however, their performance was similar among patients with systemic autoimmune diseases. A subset of SS patients is still missed by both criteria sets.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.016 | 0.060 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.003 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".