Multicenter Delphi Exercise to Identify Important Key Items for Classifying Systemic Lupus Erythematosus
Bibliographic record
Abstract
OBJECTIVE: The American College of Rheumatology and the European League Against Rheumatism embarked on a project to reevaluate classification criteria for systemic lupus erythematosus (SLE). The first phase of the classification project involved generation of a broad set of items potentially useful for classification of SLE and their selection for use in a subsequent forced-choice decision analysis. METHODS: A large international group of expert lupus clinicians was invited to participate in a 2-step process to generate, rate, and select items based on their importance in diagnosing early and established SLE, via a web-based survey. RESULTS: A total of 135 and 147 experts were invited to participate in the item-generation and item-reduction process, respectively. Of 145 items generated, item reduction resulted in 40 candidate items moving forward to the next phase. Key features for classifying both early and established SLE included characteristic autoantibodies, specific renal features, and skin manifestations. A small majority (51%) stated that 1 organ system would be sufficient for classifying SLE, but that additional typical laboratory features (antinuclear antibody, anti-double-stranded DNA) would be required. Notably, 85% of the expert group would positively classify SLE if renal pathology alone showed lupus nephritis. CONCLUSION: The Delphi exercise resulted in a set of 40 candidate criteria for the classification of SLE for subsequent assessment. This study comprised the largest panel ever involved in the development of SLE classification criteria, providing a broadly representative view of the current approach to classification of SLE.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.131 | 0.126 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.003 | 0.002 |
| Science and technology studies | 0.003 | 0.002 |
| Scholarly communication | 0.002 | 0.002 |
| Open science | 0.002 | 0.008 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.006 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".