Multicenter Delphi Exercise to Identify Important Key Items for Classifying Systemic Lupus Erythematosus
Bibliographic record
Abstract
OBJECTIVE: The American College of Rheumatology and the European League Against Rheumatism embarked on a project to reevaluate classification criteria for systemic lupus erythematosus (SLE). The first phase of the classification project involved generation of a broad set of items potentially useful for classification of SLE and their selection for use in a subsequent forced-choice decision analysis. METHODS: A large international group of expert lupus clinicians was invited to participate in a 2-step process to generate, rate, and select items based on their importance in diagnosing early and established SLE, via a web-based survey. RESULTS: A total of 135 and 147 experts were invited to participate in the item-generation and item-reduction process, respectively. Of 145 items generated, item reduction resulted in 40 candidate items moving forward to the next phase. Key features for classifying both early and established SLE included characteristic autoantibodies, specific renal features, and skin manifestations. A small majority (51%) stated that 1 organ system would be sufficient for classifying SLE, but that additional typical laboratory features (antinuclear antibody, anti-double-stranded DNA) would be required. Notably, 85% of the expert group would positively classify SLE if renal pathology alone showed lupus nephritis. CONCLUSION: The Delphi exercise resulted in a set of 40 candidate criteria for the classification of SLE for subsequent assessment. This study comprised the largest panel ever involved in the development of SLE classification criteria, providing a broadly representative view of the current approach to classification of SLE.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.003 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.002 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".