MACHINE LEARNING CAN IDENTIFY AN ANTINUCLEAR ANTIBODY PATTERN THAT MAY RULE OUT SYSTEMIC AUTOIMMUNE RHEUMATIC DISEASES
Bibliographic record
Abstract
O032 / #273 Topic:AS23 - SLE-Diagnosis, Manifestations, & Outcomes ABSTRACT CONCURRENT SESSION 05: EMERGING INSIGHTS ON THE MANAGEMENT OF LUPUS MANIFESTATIONS AND COMORBIDITIES 23-05-2025 1:40 PM - 2:40 PM Background/Purpose Antinuclear antibody (ANA) testing is used to screen for systemic autoimmune rheumatic diseases (SARD) like systemic lupus erythematosus. It is well established that a nuclear dense fine-speckled (DFS) ANA pattern (AC-2), being rare among SARD patients, decreases the likelihood of these conditions. However, the AC-2 pattern is challenging for lab technologists to accurately identify due to similarities with other patterns, ie, AC-4 (speckled) and AC-30 (nuclear speckled with mitotic plate staining), whichareassociated with SARDs. We determined if machine learning could accurately differentiate between AC-2 and SARD-related AC-4/AC-30 patterns. Methods 13,671 ANA images from SLE patients enrolled in the Systemic Lupus International Collaborating Clinics Inception Cohort (SLICC, n=2,825 images), non-SLE subjects enrolled in the Ontario Health Study (OHS, n=10,639 images), and the International Consensus on ANA Patterns (ICAP, n=207 images) were analyzed. All SLICC and OHS ANA were performed in one central laboratory using IFA on HEp-2 cells (NovaLite, Werfen, SD) and read on a digital IFA microscope (NovaView, Werfen, SD). A lab technologist (HH) with >30 years of experience identified AC-2, AC-4, and AC-30 images. Images were resized to 224x224 pixels. Three machine learning models (ANA Reader©) using a convolutional neural network (CNN) and an image feature extractor were developed to differentiate AC-2 from the other patterns. We also merged the outputs of all 3 CNNs to create a combined ANA Reader© model. 80% of the images were used for training and 20% for validation. We compared the performance of the 4 machine learning models (lab technologist as the reference standard) to determine the best prediction model. Results The lab technologist identified 308 AC-2, 957 AC-4, and 379 AC-30 images. All 4 models performed similarly with high area-under-the-curve (AUC) scores ranging from 96.5%-97.1% (Table 1). When comparing other performance metrics, the combined ANA Reader© model performed the best with the highest accuracy (93.0%), precision (92.7%), specificity (93.2%), and F1 score (92.7%). It was tied with another CNN model (Model 2) for the second most sensitive model (92.7%). Table 1. Comparison of different ANA Reader© convolutional neural network (CNN) models and a combined model to differentiate between AC-2 vs. AC-4 and AC-30 antinuclear antibody (ANA) patterns. Conclusions We developed a highly precise and accurate machine learning model, ANA Reader©, that discriminates the nuclear DFS pattern (AC-2) from other similar ANA patterns, potentially speeding up the differentiation of those at risk vs. not at risk of SARDs and reducing the need for unnecessary rheumatologic investigations or assessments. External validation of our model in other cohorts will be done before this model is adopted into laboratories and clinical practice.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.005 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.001 | 0.000 |
| Insufficient payload (model declined to judge) | 0.003 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".