Reliability assessment of ultrasound muscle echogenicity in patients with rheumatic diseases: Results of a multicenter international web-based study
Bibliographic record
Abstract
Objectives: To investigate the inter/intra-reliability of ultrasound (US) muscle echogenicity in patients with rheumatic diseases. Methods: Forty-two rheumatologists and 2 radiologists from 13 countries were asked to assess US muscle echogenicity of quadriceps muscle in 80 static images and 20 clips from 64 patients with different rheumatic diseases and 8 healthy subjects. Two visual scales were evaluated, a visual semi-quantitative scale (0-3) and a continuous quantitative measurement ("VAS echogenicity," 0-100). The same assessment was repeated to calculate intra-observer reliability. US muscle echogenicity was also calculated by an independent research assistant using a software for the analysis of scientific images (ImageJ). Inter and intra reliabilities were assessed by means of prevalence-adjusted bias-adjusted Kappa (PABAK), intraclass correlation coefficient (ICC) and correlations through Kendall's Tau and Pearson's Rho coefficients. Results: The semi-quantitative scale showed a moderate inter-reliability [PABAK = 0.58 (0.57-0.59)] and a substantial intra-reliability [PABAK = 0.71 (0.68-0.73)]. The lowest inter and intra-reliability results were obtained for the intermediate grades (i.e., grade 1 and 2) of the semi-quantitative scale. "VAS echogenicity" showed a high reliability both in the inter-observer [ICC = 0.80 (0.75-0.85)] and intra-observer [ICC = 0.88 (0.88-0.89)] evaluations. A substantial association was found between the participants assessment of the semi-quantitative scale and "VAS echogenicity" [ICC = 0.52 (0.50-0.54)]. The correlation between these two visual scales and ImageJ analysis was high (tau = 0.76 and rho = 0.89, respectively). Conclusion: The results of this large, multicenter study highlighted the overall good inter and intra-reliability of the US assessment of muscle echogenicity in patients with different rheumatic diseases.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".