Utility of the American-European Consensus Group and American College of Rheumatology Classification Criteria for Sjogren's syndrome in patients with systemic autoimmune diseases in the clinical setting
Bibliographic record
Abstract
OBJECTIVE: The aim of this study was to evaluate the feasibility and performance of the American-European Consensus Group (AECG) and ACR Classification Criteria for SS in patients with systemic autoimmune diseases. METHODS: Three hundred and fifty patients with primary SS, SLE, RA or scleroderma were randomly selected from our patient registry. Each patient was clinically diagnosed as probable/definitive SS or non-SS following a standardized evaluation including clinical symptoms and manifestations, confirmatory tests, fluorescein staining test, autoantibodies, lip biopsy and medical chart review. Using the clinical diagnosis as the gold standard, the degree of agreement with each criteria set and between the criteria sets was estimated. RESULTS: One hundred fifty-four (44%) patients were diagnosed with SS. The AECG criteria were incomplete in 36 patients (10.3%) and the ACR criteria in 96 (27.4%; P < 0.001). Nevertheless, their ability to classify patients was almost identical, with a sensitivity of 61.6 vs 62.3 and a specificity of 94.3 vs 91.3, respectively. Either set of criteria was met by 123 patients (80%); 95 (61.7%) met the AECG criteria and 96 (62.3%) met the ACR criteria, but only 68 (44.2%) patients met both sets. The concordance rate between clinical diagnosis and AECG or ACR criteria was moderate (k statistic 0.58 and 0.55, respectively). Among 99 patients with definitive SS sensitivity was 83.3 vs 77.7 and specificity was 90.8 vs 85.6, respectively. A discrepancy between clinical diagnosis and criteria was seen in 59 patients (17%). CONCLUSION: The feasibility of the SS AECG criteria is superior to that of the ACR criteria, however, their performance was similar among patients with systemic autoimmune diseases. A subset of SS patients is still missed by both criteria sets.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".