International multicenter psoriasis and psoriatic arthritis reliability trial for the assessment of skin, joints, nails, and dactylitis
Bibliographic record
Abstract
OBJECTIVE: Clinical trials in psoriasis and psoriatic arthritis (PsA) involve assessment of the skin and joints. This study aimed to determine whether assessment of the skin and joints in patients with PsA by rheumatologists and dermatologists is reproducible. METHODS: Ten rheumatologists and 9 dermatologists from 7 countries met for a combined physical examination exercise to assess 20 PsA patients (11 men, mean age 51 years, mean PsA duration 11 years). Each physician assessed 10 patients according to a modified Latin square design that enabled the assessment of patient, assessor, and order effect. Tender joint count (TJC), swollen joint count (SJC), dactylitis, physician's global assessment (PGA) of PsA disease activity (PGA-PsA), psoriasis body surface area (BSA), Psoriasis Area and Severity Index (PASI), Lattice System Physician's Global Assessment of psoriasis (LS-PGA), National Psoriasis Foundation Psoriasis Score (NPF-PS), modified Nail Psoriasis Severity Index (mNAPSI), number of fingernails with nail changes (NN), and PGA of psoriasis activity (PGA-Ps) were assessed. Variance components analyses were carried out to estimate the intraclass correlation coefficient (ICC), adjusted for the order of measurements. RESULTS: There is excellent agreement (ICC >/=0.80) on the mNAPSI, substantial agreement (0.6 >/= ICC < 0.80) on the TJC, PASI, and NN, moderate agreement (0.4 >/= ICC < 0.60) on the PGA-Ps, LS-PGA, NPF-PS, and BSA, and fair agreement (0.2 >/= ICC < 0.40) on the SJC, dactylitis, and PGA-PsA. The only measure that showed a significant difference between dermatologists and rheumatologists was dactylitis (P = 0.0005). CONCLUSION: There is substantial to excellent agreement on the TJC, PASI, NN, and mNAPSI among rheumatologists and dermatologists.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".