OP0047 Estimating the impact of a training workshop on the reliability of salivary gland ultrasound scoring in Sjögren's syndrome
Bibliographic record
Abstract
<h2>Abstract</h2><h3>Background:</h3> In the recent years, salivary gland ultrasound (SGUS) has emerged as a promising diagnostic tool for SjD. SGUS is a non-invasive, cheap, and widely accessible method for the assessment of glandular abnormalities, offering high sensitivity and specificity [1-3]. However, the reliability of SGUS depends heavily on operator proficiency [4]. Standardized training could be a potential approach to improve SGUS reliability and enhance the ability of clinicians to assess the salivary glands, especially for observers with limited experience. <h3>Objectives:</h3> The aim of this study was to assess the impact of an international training workshop on inter-observer reliability of SGUS scoring in SjD. <h3>Methods:</h3> 25 healthcare professionals from 10 countries with different levels of SGUS expertise participated in the reliability exercise. Participants assessed 8 SGUS images of 20 patients suspected of SjD, before and after attending a training workshop. The images consisted of greyscale (GS) and color Doppler (CD) scans of the submandibular and parotid glands. Participants scored the images using the OMERACT GS and CD scoring systems. Intraclass correlation coefficients (ICC) were used to assess overall inter-observer reliability and ICCs of participants vs SGUS experts (expert-participant reliability) before and after the workshop. Analyses were stratified according to SGUS experience, i.e. participants without experience and participants with ≥1 year experience. <h3>Results:</h3> The overall pre-workshop inter-observer reliability ICC for the total OMERACT score was 0.68 for GS and 0.73 for CD. Post-workshop overall inter-observer reliability was 0.79 for GS and 0.72 for the CD. The training led to significant improvements for the total OMERACT GS score, with an increase in ICC by 0.06±0.12 (p=0.020), when comparing participants to experts. Most improvement was found in the inter-observer reliability of the submandibular glands. For CD, the total OMERACT score showed a non-significant ICC improvement of 0.03±0.09 (p=0.129). Participants without prior SGUS experience demonstrated significant improvement in the total OMERACT GS score, with ICC increasing by 0.13±0.13 (p=0.012), compared to a negligible improvement of 0.01±0.09 (p=0.624) among the experienced group. We found no significant differences in the impact of the workshop on the evaluation of CD images between the compared groups. <h3>Conclusion:</h3> The international training workshop significantly improved the reliability in scoring greyscale SGUS, particularly when assessing the submandibular glands and for participants without prior experience. Effects were limited in CD scoring and in participants with over one year of SGUS experience. <h3>REFERENCES:</h3> [1] Ramsubeik K, Motilal S, Sanchez-Ramos L, Ramrattan LA, Kaeley GS, Singh JA. Diagnostic accuracy of salivary gland ultrasound in Sjögren's syndrome: A systematic review and meta-analysis. <i>Ther Adv Musculoskelet Dis</i>. 2020;12:1759720X20973560. Published 2020 Nov 21. doi:10.1177/1759720X20973560. [2] Jousse-Joulin S, Gatineau F, Baldini C, et al. Weight of salivary gland ultrasonography compared to other items of the 2016 ACR/EULAR classification criteria for Primary Sjögren's syndrome. <i>J Intern Med</i>. 2020;287(2):180-188. doi:10.1111/joim.12992. [3] Hocevar A, Ambrozic A, Rozman B, Kveder T, Tomsic M. Ultrasonographic changes of major salivary glands in primary Sjogren's syndrome. Diagnostic value of a novel scoring system. Rheumatology (Oxford). 2005;44(6):768-772. doi:10.1093/rheumatology/keh588. [4] Saied F, Włodkowska-Korytkowska M, Maślińska M, et al. The usefulness of ultrasound in the diagnostics of Sjögren's syndrome. <i>J Ultrason</i>. 2013;13(53):202-211. doi:10.15557/JoU.2013.0020. <h3>Acknowledgements:</h3> <b>NIL</b>. <h3>Disclosure of Interests:</h3> <b>None declared</b>. © The Authors 2025. This abstract is an open access article published in Annals of Rheumatic Diseases under the CC BY-NC-ND license (http://creativecommons.org/licenses/by-nc-nd/4.0/). Neither EULAR nor the publisher make any representation as to the accuracy of the content. The authors are solely responsible for the content in their abstract including accuracy of the facts, statements, results, conclusion, citing resources etc.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.005 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".