SLE-DISEASOME: A WIDE SPECTRUM DATABASE OF SYSTEMIC LUPUS ERYTHEMATOSUS RELEVANT FUNCTIONAL PATHWAYS
Bibliographic record
Abstract
PT017 / #649 Topic: AS12 - Genetics, Epigenetics, Transcriptomics POSTER TOUR 04: SLE PATHOGENESIS 24-05-2025 10:00 AM - 10:20 AM Background/Purpose Systemic Lupus Erythematosus (SLE) is an autoimmune disease characterized by unpredictable patterns of flares and remissions, affecting a wide range of tissues and organs. SLE causes significant suffering and mortality, and treatment efficacy varies enormously among patients. The main contributing factor to treatment failure and the broad clinical spectrum observed is the heterogeneity in dysregulated molecular mechanisms across patients. Therefore, the use of personalized therapies based on molecular information is considered a promising strategy to address disease heterogeneity, although their practical implementation still faces substantial challenges.[1] Transcriptomics offers a powerful tool for understanding molecular profiles, but its reproducibility is strongly affected by batch effects across studies, and its high dimensionality makes interpretation difficult for clinical practice. Pathway-based single-sample scoring approaches emerge as a potential solution by translating gene expression into standardized activity measurements using small sets of functional pathways.[2] The key step is to define the disease-relevant biological pathways. There are numerous databases of biological functions available. However, using a single pathway database may lead to biased results based on the knowledge collected in that particular database, while using multiple databases could result in redundant findings. Additionally, selecting significant pathways using 1 specific study can yield cohort-dependent results, many of which may not be reproducible in other studies. Therefore, in this study, we defined a comprehensive collection of disease-relevant gene-signatures, called the SLE-diseasome, based on a multicohort approach and integrating multiple layers of database-derived biological knowledges. Methods For the development of the SLE-diseasome (Figure 1), a total of 16 SLE datasets, comprising about 5500 SLE patient data and 900 healthy samples were used as well as 11 different pathway databases. The different pathways were divided into subpathways, or gene-signatures, using a co-occurrence-based k-means clustering across datasets, to get molecular and functional granularity. Redundancy across all these gene-signatures was reduced by filtering pathways based on similarity, using the Jaccard index. Upon each step, the pathway database was re-annotated. Next, each pathway and patient was scored from each study using m-score-based single-sample molecular scoring. Significance with respect to healthy distribution at patient and pathway level was also calculated and incorporated. Disease-relevant pathways were defined as pathways that were significant in at least 10% of the patients when compared to healthy controls, and significant across 7 different studies. These 2 parameters were internally optimized to keep the data structure and minimizing false positive results. Significant pathways were clustered and re-annotated. Figure 1: SLE-Diseasome database development workflow. Results We obtained a total of 4400 SLE-relevant and robust functional pathways integrating 16 SLE datasets and 11 different pathway databases. By obtaining clusters of pathways from different initial sources, we can go 1 step further when interpreting results, establishing connections between different functions and annotations. The applicability of the SLE-diseasome was tested in different scenarios, for patient stratification analysis and for the generation and cross-cohort validation of machine learning models to predict clinical manifestations and drug response. Conclusions The SLE-diseasome offers a new SLE-specific database connecting multiple layers of database-derived biological knowledge. It is defined using a robust multicohort approach, bringing us 1 step closer to the effective use of molecular information in clinical practice through single-sample molecular scoring. References: [1.] Toro-Dominguez D. Brief Bioinform 2022;23(5):bbac332. [2.] Foroutan M. BMC Bioinformatics 2018;19:404.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.005 |
| Meta-epidemiology (narrow) | 0.002 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.008 | 0.006 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.003 | 0.002 |
| Open science | 0.002 | 0.004 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.010 | 0.005 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".