Reliability assessment of the definition of ultrasound enthesitis in SpA: results of a large, multicentre, international, web-based study
Bibliographic record
Abstract
OBJECTIVES: To investigate the reliability of the OMERACT US Task Force definition of US enthesitis in SpA. METHODS: In this web exercise, based on the evaluation of 101 images and 39 clips of the main entheses of the lower limbs, the elementary components included in the OMERACT definition of US enthesitis in SpA (hypoechoic areas, entheseal thickening, power Doppler signal at the enthesis, enthesophytes/calcifications, bone erosions) were assessed by 47 rheumatologists from 37 rheumatology centres in 15 countries. Inter- and intra-observer reliability of the US components of enthesitis was calculated using Light's kappa, Cohen's kappa, Prevalence And Bias Adjusted Kappa (PABAK) and their 95% CIs. RESULTS: Bone erosions and power Doppler signal at the enthesis showed the highest overall inter-reliability [Light's kappa: 0.77 (0.76-0.78), 0.72 (0.71-0.73), respectively; PABAK: 0.86 (0.86-0.87), 0.73 (0.73-0.74), respectively], followed by enthesophytes/calcifications [Light's kappa: 0.65 (0.64-0.65), PABAK: 0.67 (0.67-0.68)]. This was moderate for entheseal thickening [Light's kappa: 0.41 (0.41-0.42), PABAK: 0.41 (0.40-0.42)], and fair for hypoechoic areas [Light's kappa: 0.37 (0.36-0.38); PABAK: 0.37 (0.37-0.38)]. A similar trend was observed in the intra-reliability exercise, although this was characterized by an overall higher degree of reliability for all US elementary components compared with the inter-observer evaluation. CONCLUSIONS: The results of this multicentre, international, web-based study show a good reliability of the OMERACT US definition of bone erosions, power Doppler signal at the enthesis and enthesophytes/calcifications. The low reliability of entheseal thickening and hypoechoic areas raises questions about the opportunity to revise the definition of these two major components for the US diagnosis of enthesitis.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.031 | 0.063 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.001 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".