MétaCan
Menu
Back to cohort
Record W4416296293 · doi:10.1142/s2661341725740438

Deep Learning on Spinal SPARCC Scoring System in Axial SpA Cohort

2025· article· en· W4416296293 on OpenAlexaboutno aff
Eugenia Yue Tian Lung, Yingying Lin, Ho Yin Chung, Peng Cao, Kam Ho Lee, Vince Wing Hang Lau, Brian Yu, Weiqiang Lin, Shirley Chiu Wai Chan

Bibliographic record

VenueJournal of Clinical Rheumatology and Immunology · 2025
Typearticle
Languageen
FieldMedicine
TopicSpondyloarthritis Studies and Treatments
Canadian institutionsnot available
Fundersnot available
KeywordsIntraclass correlationMagnetic resonance imagingPearson product-moment correlation coefficientDeep learningAxial spondyloarthritisCohortSørensen–Dice coefficient

Abstract

fetched live from OpenAlex

Background: Deep learning models based on the Spondyloarthritis Research Consortium of Canada (SPARCC) scoring system have been used in previous studies to assess sacroiliac joint inflammation in patients with axial spondyloarthritis (SpA). However, these patients also commonly have active spinal inflammation, and detecting these changes holds diagnostic, prognostic, and therapeutic significance. This study aimed to develop a deep learning algorithm for spinal inflammation based on SPARCC scoring system in patients with axial SpA. Methods: The study cohort included 330 participants with axial SpA. All patients underwent whole spine magnetic resonance imaging (MRI) with short tau inversion recovery (STIR) sequence by 3T MR unit. Three independent readers identified regions of interest (ROIs) to identify bone marrow edema (BME) and performed SPARCC scoring. Two deep learning models based on attention Unet were trained. The BME model was employed to differentiate image with or without spinal inflammation and to delineate BMEs. The vertebral body-intervertebral disc model was utilized to identify discovertebral units. Setting the threshold brightness was applied to detect the region with CSF. The intraclass correlation coefficient (ICC) and Pearson coefficient were used to evaluate the agreement and the correlation between the score of human readers and the score of deep learning-based pipeline. Performance of the models was evaluated using sensitivity, specificity, accuracy, and the Dice coefficient. Results: The ICC and the Pearson coefficient between the SPARCC scores from three human readers and the deep learning-based scoring pipeline were 0.80 and 0.82, respectively. The sensitivity and specificity of identifying image with spinal inflammation were 0.90 and 0.84, respectively. The accuracy of identifying the region containing CSF was 0.88 in images with spinal inflammation. The Dice coefficients were 0.81 (vertebral bodies) and 0.80 (intervertebral disc) in images with spinal inflammation. Conclusion: The high consistency with human readers suggested that the deep learning-based pipeline could offer a SPARCC-informed approach for scoring spinal STIR images in axial SpA.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.001
metaresearch head score (Gemma)0.001
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: Observational
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.040
Threshold uncertainty score0.471

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0010.001
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0010.000
Bibliometrics0.0000.000
Science and technology studies0.0000.000
Scholarly communication0.0000.000
Open science0.0000.000
Research integrity0.0000.001
Insufficient payload (model declined to judge)0.0000.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.024
GPT teacher head0.361
Teacher spread0.338 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designObservational
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2025
Admission routes1
Has abstractyes

Explore more

Same venueJournal of Clinical Rheumatology and ImmunologySame topicSpondyloarthritis Studies and TreatmentsFrench-language works237,207