Items for developing revised classification criteria in systemic sclerosis: Results of a consensus exercise
Bibliographic record
Abstract
OBJECTIVE: Classification criteria for systemic sclerosis (SSc; scleroderma) are being updated. Our objective was to select a set of items potentially useful for the classification of SSc using consensus procedures, including the Delphi and nominal group techniques (NGT). METHODS: Items were identified through 2 independent consensus exercises performed by the Scleroderma Clinical Trials Consortium and the European League Against Rheumatism Scleroderma Trials and Research Group. The first-round items from both exercises were collated and redundancies were removed, leaving 168 items. A 3-round Delphi exercise was performed using a 1-9 scale (where 1 = completely inappropriate and 9 = completely appropriate) and a consensus meeting using NGT was conducted. During the last Delphi round, the items were ranked on a 1-10 scale. RESULTS: In round 1, 106 experts rated the 168 items. Those with a median score of <4 were removed, resulting in a list of 102 items. In round 2, the items were again rated for appropriateness and subjected to a consensus meeting using NGT by European and North American SSc experts (n = 16), resulting in 23 items. In round 3, SSc experts (n = 26) then individually scored each of the 23 items in a last Delphi round using an appropriateness score (1-9) and ranking their 10 most appropriate items for the classification of SSc. Presence of skin thickening, SSc-specific autoantibodies, abnormal nailfold capillary pattern, and Raynaud's phenomenon ranked highest in the final list that also included items indicating internal organ involvement. CONCLUSION: The Delphi exercise and NGT resulted in a set of 23 items for the classification of SSc that will be assessed for their discriminative properties in a prospective study.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".