International spondyloarthritis interobserver reliability exercise--the INSPIRE study: II. Assessment of peripheral joints, enthesitis, and dactylitis.
Bibliographic record
Abstract
OBJECTIVE: To determine whether the assessments of peripheral joints and enthesitis were reproducible for both AS and PsA with axial disease, and whether dactylitis assessment is reproducible in patients with PsA. METHODS: A group of 20 rheumatologists from 11 countries with expertise in spondyloarthritis (SpA) met for a combined physical examination exercise to assess 10 patients with PsA with axial involvement (9 men, 1 woman, mean age 52 yrs, disease duration 17 yrs) and 9 patients with AS (7 men, 2 women, mean age 38 yrs, disease duration 16 yrs). A modified Latin-square design that enabled assessment of patient, assessor, and order effect was used. Measures included were number of tender and swollen joints, presence of enthesitis using 6 different indices, and dactylitis score. Data were analyzed using intraclass correlation (ICC) adjusted for order of measurements. RESULTS: The majority of the variance was contributed by the patients. There was no order effect. The assessment of tender joints (ICC 0.69) was more reliable than the assessment of swollen joints (ICC 0.54). Moreover, there was better agreement in patients with PsA (ICC 0.78) than in patients with AS (ICC 0.62). There was excellent agreement on the number of active enthesitis sites (ICC 0.86). All the enthesitis indices provided substantial to excellent agreement among observers. Agreement for the dactylitis score was substantial (ICC 0.70). CONCLUSION: The assessment of peripheral joints is more reliable in patients with PsA. Enthesitis instruments can be used reliably in patients with AS and patients with PsA with spinal involvement. The Leeds dactylitis instrument functions well in PsA.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".