Items for developing revised classification criteria in systemic sclerosis: Results of a consensus exercise
Bibliographic record
Abstract
OBJECTIVE: Classification criteria for systemic sclerosis (SSc; scleroderma) are being updated. Our objective was to select a set of items potentially useful for the classification of SSc using consensus procedures, including the Delphi and nominal group techniques (NGT). METHODS: Items were identified through 2 independent consensus exercises performed by the Scleroderma Clinical Trials Consortium and the European League Against Rheumatism Scleroderma Trials and Research Group. The first-round items from both exercises were collated and redundancies were removed, leaving 168 items. A 3-round Delphi exercise was performed using a 1-9 scale (where 1 = completely inappropriate and 9 = completely appropriate) and a consensus meeting using NGT was conducted. During the last Delphi round, the items were ranked on a 1-10 scale. RESULTS: In round 1, 106 experts rated the 168 items. Those with a median score of <4 were removed, resulting in a list of 102 items. In round 2, the items were again rated for appropriateness and subjected to a consensus meeting using NGT by European and North American SSc experts (n = 16), resulting in 23 items. In round 3, SSc experts (n = 26) then individually scored each of the 23 items in a last Delphi round using an appropriateness score (1-9) and ranking their 10 most appropriate items for the classification of SSc. Presence of skin thickening, SSc-specific autoantibodies, abnormal nailfold capillary pattern, and Raynaud's phenomenon ranked highest in the final list that also included items indicating internal organ involvement. CONCLUSION: The Delphi exercise and NGT resulted in a set of 23 items for the classification of SSc that will be assessed for their discriminative properties in a prospective study.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.239 | 0.312 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.003 | 0.004 |
| Bibliometrics | 0.009 | 0.006 |
| Science and technology studies | 0.002 | 0.002 |
| Scholarly communication | 0.004 | 0.004 |
| Open science | 0.004 | 0.006 |
| Research integrity | 0.002 | 0.004 |
| Insufficient payload (model declined to judge) | 0.004 | 0.003 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".