MULTI-OMIC INTEGRATION REVEALS THREE MOLECULAR SUBTYPES WITH DISTINCT IMMUNOLOGICAL PHENOTYPES IN A COHORT OF 722 SYSTEMIC LUPUS ERYTHEMATOSUS PATIENTS
Bibliographic record
Abstract
O026 / #740 Topic: AS22 - SLE Heterogeneity Late-Breaking Abstract ABSTRACT CONCURRENT SESSION 04: ADVANCING LUPUS THERAPIES AND INSIGHTS 22-05-2025 1:40 PM - 2:40 PM Background/Purpose Integrative omics approaches offer a powerful strategy to dissect the complex biological networks and pathways involved in disease pathophysiology. However, integrating diverse data types with varying scales, biological contexts, and feature numbers poses significant challenges. This study aims to: i) identify molecular subtypes of SLE patients via a multi-omic integrative approach, ii) characterize these clusters using molecular and clinical data, and iii) identify the most discriminant subset of features for patient classification. Methods Similarity Network Fusion (SNF) was used to integrate baseline transcriptomic (RNAseq), proteomic (Olink), and epigenomic (EMseq) data from the whole blood of 722 SLE patients that were randomized to placebo plus standard of care in phase 3 clinical trials ( NCT03616964 , NCT03616912 ), and 84 healthy controls. Patient subgroups were identified via spectral clustering, and cluster robustness was determine using a bootstrapping approach (n = 30). Clusters were characterized with omics data via differential expression and gene set enrichment analysis (GSEA). Clusters were also characterized with clinical metrics via t-tests and random forest analysis. Finally, we utilized Data Integration Analysis for Biomarker Discovery using Latent cOmponents (DIABLO) modeling to identify potential biomarkers associated with these SLE classes. Results Integration of the omics datatypes identified 3 distinct clusters of individuals: cluster 1 (n = 176), cluster 2 (n = 299), and cluster 3 (n = 331). All 84 healthy controls were grouped within cluster 2 (Figure 1A). Notably, clustering based on individual datatypes failed to reproduce these distinct clusters. Clinical data revealed that SLE patients in cluster 2 exhibited significantly lower dsDNA, IFI44L, and SLEDAI scores, along with higher complement (C3) levels compared to the other clusters (Figure 1B), indicating a milder form of SLE. This is consistent with the fact that these SLE patients clustered with the healthy controls. Additional clinical measurements (Figure 1C) and pathway enrichment analysis (Figure 1D) revealed distinct signatures in the other 2 clusters: cluster 3 displayed an elevated adaptive immunity signature, while cluster 1 was characterized by innate immunity signatures. Finally, integrating all 3 datatypes using the DIABLO modeling approach identified 3 latent components with a minimal set of discriminating biomarkers for each cluster. A) Spectral clustering of patient-patient similarities calculated from the SNF-fused data. B) Distribution of conventional SLE metrics among the 3 SNF clusters (excluding healthy patients). Significant differences assessed with a t-test and false discover rate (FDR). C) Significant clinical differences between Cluster 1 and Cluster 3 using t-test and FDR. Fold change (C1/C3) greater than 1 indicates increased measurements in Custer 1 compared to Cluster 3. D) Normalized Enrichment Score (NES) of select GO terms from a GSEA analysis comparing C1 vs C3 using RNAseq. Positive NES indicates increased activity in C1, and vice-versa. Figure 1. SLE patient clusters and characterization. Conclusions Multi-omic integration revealed 3 molecularly distinct clusters of SLE patients. Using orthogonal datatypes (clinical and omics), we characterized these clusters into 3 novel classifications: mild, innate-driven, and adaptive-driven immunity. This work helps our understanding of the complex heterogenous nature of SLE and will guide targeted treatment approaches with innate or adaptive involvement.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".