Joint count reliability in psoriatic arthritis observational trials--an unreported problem
Bibliographic record
Abstract
Sir, Multiple observers are a reality of large observational and multicentre studies and introduce the challenge of addressing inter-rater reliability. Long term Outcomes in Psoriatic Arthritis II (LOPAS II) is a multicentre prospective observational study investigating work disability in PsA. The primary endpoint is presenteeism (reduced effectiveness at work), but the secondary endpoints include tender and swollen joint scores. Clinical assessments will be undertaken at multiple sites across the UK. We set out to undertake a reliability exercise to estimate joint count reliability in LOPAS II. We invited assessors from each centre to an education day at the lead site. A 1-h seminar on the study was followed by a 45-min clinical training session on joint counts lead by two trainers, each with >10 years experience in PsA joint assessment. The session was followed by a joint count reliability exercise. Four patients of differing disease duration (1–33 years) and activity (from 22 tender and 9 swollen to 2 tender and 2 swollen joints as assessed by the instructors) were assessed using a modified (asymmetrical) Latin square design. Reliability was measured using Krippendorff’s α, a reliability coefficient that accommodates the modified Latin square design [1]. Analyses were undertaken on the group as a whole and then repeated excluding those who self-reported to be unconfident or who had never performed joint assessments before. Twelve assessors from seven units attended: one doctor, seven rheumatology nurse specialists, one occupational therapist and three research nurses (of whom one had rheumatology experience). Reliability is reported in Table 1. Inter-rater reliability for all was low irrespective of experience, but was higher among those with experience. Joint count reliability using Krippendorff’s α Joint count reliability using Krippendorff’s α There are limited reports of joint count reliability among physicians with an interest in PsA [2–5]. Even among such experts, inter-class correlation coefficients are poor for determining peripheral joint swelling—0.13, 0.55, 0.242—and moderate for tenderness (or activity)—0.73, 0.75, 0.72. It is noteworthy that none of these studies has included the wider multidisciplinary team. To our knowledge, none of the recently published large observational studies or registry reports has reported on joint count reliability [6–10]. Only the Toronto research group has reported on joint count reliability, and this was at the time of the cohort’s inception [4]. The Toronto study involved three rheumatologists and two trainees assessing five patients in a Latin square design. There was a <1% observer variance, indicating good reliability of assessment. To enable direct comparison, the analysis of variance in our study showed that the proportion of variance attributable to (all) raters was much higher; swollen joints 56% (P = 0.094) and tender joints 60% (P = 0.004). It is noteworthy that the Toronto study was undertaken over 20 years ago and since that time the expansion of the multidisciplinary team has meant that clinical assessments are now performed by a wider clinical team including doctors, nurses and extended scope therapists. The general lack of reporting of joint count reliability may reflect a mixture of publication bias, insufficient recognition of the potential problem or misplaced confidence. The poor reliability identified in our study is important to our current study (LOPAS II) and also to assessors from centres in the UK and further afield who collect data for other large observational studies in PsA. The joint count training offered in the LOPAS II training day was minimal, as we had only anticipated the need for some fine tuning to standardize the assessment techniques. More training is required, as was mentioned by the assessors themselves in the feedback from the training day. We are attempting to standardize assessments and improve our reliability by using an instructional joint count training video as well as offering one-to-one tuition at the lead site. We are also encouraging a period of mentoring within each unit for those with less experience as well as aiming to use the same assessor to perform the joint counts at serial appointments. A repeat assessment day is planned once all centres have completed the training. To our knowledge, this is the first study investigating the joint count reliability among assessors routinely contributing data in the PsA clinical and research setting. We suggest that future reporting of joint count outcomes should include some assessment of joint count reliability in order to interpret results, particularly negative findings. Furthermore, we suggest that to optimize data collection, individual units document joint count reliability with a view to determining a potential training need. We would like to thank Nina Griffith, Nicola Waldron, Sarah Whitford, Karen Brown, Rebecca Rowland, Jennifer Brown, Beverly Vale, Julie Taylor, Wendy Wilmott, Miriam Skelton, Mandy Knight and Charlotte Cavill for their involvement with the training day. Funding: This work was supported by the National Institute of Health Research (NIHR) through the Comprehensive Clinical Research Network (to W.T.) and an unrestricted grant form Abbott Laboratories Ltd. Disclosure statement: The authors have declared no conflicts of interest.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.765 | 0.868 |
| Meta-epidemiology (narrow) | 0.002 | 0.003 |
| Meta-epidemiology (broad) | 0.008 | 0.007 |
| Bibliometrics | 0.005 | 0.010 |
| Science and technology studies | 0.002 | 0.011 |
| Scholarly communication | 0.007 | 0.007 |
| Open science | 0.007 | 0.006 |
| Research integrity | 0.007 | 0.010 |
| Insufficient payload (model declined to judge) | 0.003 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".