MétaCan
Menu
← Back to cohort
Record W4410512969 · doi:10.3899/jrheum.2025-0390.o016

RELIABILITY OF SLE-DAS, SLEDAI-2K, AND PGA INSTRUMENTS IN ASSESSING SLE DISEASE ACTIVITY: A STUDY AMONG GLOBAL LUPUS EXPERTS

2025· article· en· W4410512969 on OpenAlexaffvenue
Beatriz Mendes, Carolina Mazeda, Diogo Jesús, Carla Henriques, Ana Matos, Irene Altabás González, Samuel Lacerda de Andrade, Simone Appenzeller, George Βertsias, Ricard Cervera, Nathalie Costedoat-Chalumeau, Antonis Fanouriakis, Thiago Sotero Fragoso, Mariele Gatto, László Kovács, Chi Chiu Mok, Ioannis Parodis, José María Pego‐Reigosa, Matteo Piga, Anisur Rahman, Christopher Sjöwall, Maria G. Tektonidou, Zahi Touma, Manuel F. Ugarte‐Gil, Margherita Zen, Andrea Doria, Luís Inês

Bibliographic record

VenueThe Journal of Rheumatology · 2025
Typearticle
Languageen
FieldMedicine
TopicSystemic Lupus Erythematosus Research
Canadian institutionsToronto Western Hospital
Fundersnot available
KeywordsMedicineSystemic lupus erythematosusReliability (semiconductor)Lupus erythematosusImmunologyDiseasePhysical therapyInternal medicineAntibody

Abstract

fetched live from OpenAlex

O016 / #589 Topic: AS23 - SLE-Diagnosis, Manifestations, & Outcomes ABSTRACT CONCURRENT SESSION 02: SLE METRICS – IMPROVING OUTCOMES & MEASURES 22-05-2025 1:40 PM - 2:40 PM Background/Purpose Reliability measures the consistency of an instrument’s assessments. Instruments intended for clinical and research use must exhibit high reliability. There is a need for studies that evaluate and compare the reliability of different instruments used to measure SLE disease activity. This study aims to estimate the intrarater and interrater reliability of 3 SLE disease activity measurement tools as assessed by lupus experts: the SLE Disease Activity Score (SLE-DAS), the SLE Disease Activity Index 2000 (SLEDAI-2K), and the Physician Global Assessment (PGA).[1,2] Methods A group of 19 lupus experts from 12 countries (Europe, North America, South America, and Asia) evaluated 24 clinical case vignettes of SLE covering a wide spectrum of organ manifestations and disease severity. All raters completed a training on scoring rules for SLE-DAS, SLEDAI-2K and PGA before assessing the clinical vignettes. Raters scored each clinical vignette with SLE-DAS, SLEDAI-2K, and PGA twice, with at least a 10-day interval between rounds. The clinical vignettes were randomly ordered and assessed through an online survey. Intrarater and interrater reliability were assessed using the Intraclass Correlation Coefficient (ICC) and reported with 95% CI. For this analysis, ICC estimates were derived from a two-way random effects model (single rater). All calculations were performed using the Stata Statistical Software (Release 17) with the kappaetc module. The Coefficient of Variation (CV) was also used as a measure of reliability.[3] The CV was calculated for each vignette, based on the 19 measurements from each rater, for SLE-DAS, SLEDAI-2K, and PGA. Then, for each disease activity measure, the mean of the CV values across the 24 vignettes was used as a summary measure of within-subject variability and is expressed as a percentage. Results The 24 clinical vignettes represented a wide variety of active SLE manifestations, including skin rash (20.8%), arthritis (12.5%), renal involvement (12.5%), thrombocytopenia (12.5%), cardiac/pulmonary involvement (12.5%), mucocutaneous vasculitis (8.3%), serositis (8.3%), and neuropsychiatric SLE (8.3%). Systemic vasculitis, myositis, alopecia, hemolytic anemia, and leukopenia were each present in 4.2% of the vignettes. Hypocomplementemia and/or positive anti-dsDNA were present in 75.0%. All the 19 lupus experts completed 2 rounds of assessment of the 24 clinical vignettes, totaling 912 case assessments. Scores ranged from 0.37 to 27.37 in SLE-DAS, 0 to 21 in SLEDAI-2K, and 0.0 to 3.0 in PGA. The interrater ICCs were 0.93, 0.91, and 0.74, and the intrarater ICCs were 0.94, 0.93, and 0.88 for SLE-DAS, SLEDAI-2K, and PGA, respectively. The CVs (first rating round) were 8.2%, 19.7%, and 41.1% for SLE-DAS, SLEDAI-2K, and PGA, respectively. The ICCs (95% CI) and CVs for SLE-DAS, SLEDAI-2K, and PGA are detailed in Table 1 and Table 2. Table 1. Inter-rater and intra-rater Intraclass Correlation Coefficient (ICC) of SLE-DAS, SLEDAI-2K and PGA. Table 2. Coefficient of Variation (CV) of SLE-DAS, SLEDAI-2K and PGA. Conclusions This study demonstrates that both SLE-DAS and SLEDAI-2K presented good to excellent interrater reliability, indicating strong consistency in scoring across different experts. Notably, SLE-DAS achieved excellent intrarater reliability, reflecting a high degree of stability in individual assessments between assessments at different times. Both SLEDAI-2K and PGA exhibited good to excellent intrarater reliability, while PGA showed moderate to good interrater reliability. Furthermore, SLE-DAS exhibited the lowest within-subject variability, as evidenced by its lower CV values compared to SLEDAI-2K and PGA. References: [1.] Jesus D. Ann Rheum Dis 2019;78:365-71. [2.] Piga M. Lancet Rheumatol 2022;4:e441-9. [3.] Shechtman O. In: S.A.R. Doi, G.M. Williams (Eds.). Methods Clin Epidemiol 2013:39-49.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.047
metaresearch head score (Gemma)0.096
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch
Consensus categoriesnone
DomainCandidate signal: Methods · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: Observational
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.953
Threshold uncertainty score0.247

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0470.096
Meta-epidemiology (narrow)0.0000.001
Meta-epidemiology (broad)0.0010.001
Bibliometrics0.0030.001
Science and technology studies0.0010.002
Scholarly communication0.0010.002
Open science0.0010.002
Research integrity0.0010.001
Insufficient payload (model declined to judge)0.0010.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.018
GPT teacher head0.351
Teacher spread0.333 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

Study designObservational
DomainMethods
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations1
Published2025
Admission routes2
Has abstractyes

Explore more

Same venueThe Journal of Rheumatology→Same topicSystemic Lupus Erythematosus Research→French-language works237,207→