SLE-DISEASOME: A WIDE SPECTRUM DATABASE OF SYSTEMIC LUPUS ERYTHEMATOSUS RELEVANT FUNCTIONAL PATHWAYS
Notice bibliographique
Résumé
PT017 / #649 Topic: AS12 - Genetics, Epigenetics, Transcriptomics POSTER TOUR 04: SLE PATHOGENESIS 24-05-2025 10:00 AM - 10:20 AM Background/Purpose Systemic Lupus Erythematosus (SLE) is an autoimmune disease characterized by unpredictable patterns of flares and remissions, affecting a wide range of tissues and organs. SLE causes significant suffering and mortality, and treatment efficacy varies enormously among patients. The main contributing factor to treatment failure and the broad clinical spectrum observed is the heterogeneity in dysregulated molecular mechanisms across patients. Therefore, the use of personalized therapies based on molecular information is considered a promising strategy to address disease heterogeneity, although their practical implementation still faces substantial challenges.[1] Transcriptomics offers a powerful tool for understanding molecular profiles, but its reproducibility is strongly affected by batch effects across studies, and its high dimensionality makes interpretation difficult for clinical practice. Pathway-based single-sample scoring approaches emerge as a potential solution by translating gene expression into standardized activity measurements using small sets of functional pathways.[2] The key step is to define the disease-relevant biological pathways. There are numerous databases of biological functions available. However, using a single pathway database may lead to biased results based on the knowledge collected in that particular database, while using multiple databases could result in redundant findings. Additionally, selecting significant pathways using 1 specific study can yield cohort-dependent results, many of which may not be reproducible in other studies. Therefore, in this study, we defined a comprehensive collection of disease-relevant gene-signatures, called the SLE-diseasome, based on a multicohort approach and integrating multiple layers of database-derived biological knowledges. Methods For the development of the SLE-diseasome (Figure 1), a total of 16 SLE datasets, comprising about 5500 SLE patient data and 900 healthy samples were used as well as 11 different pathway databases. The different pathways were divided into subpathways, or gene-signatures, using a co-occurrence-based k-means clustering across datasets, to get molecular and functional granularity. Redundancy across all these gene-signatures was reduced by filtering pathways based on similarity, using the Jaccard index. Upon each step, the pathway database was re-annotated. Next, each pathway and patient was scored from each study using m-score-based single-sample molecular scoring. Significance with respect to healthy distribution at patient and pathway level was also calculated and incorporated. Disease-relevant pathways were defined as pathways that were significant in at least 10% of the patients when compared to healthy controls, and significant across 7 different studies. These 2 parameters were internally optimized to keep the data structure and minimizing false positive results. Significant pathways were clustered and re-annotated. Figure 1: SLE-Diseasome database development workflow. Results We obtained a total of 4400 SLE-relevant and robust functional pathways integrating 16 SLE datasets and 11 different pathway databases. By obtaining clusters of pathways from different initial sources, we can go 1 step further when interpreting results, establishing connections between different functions and annotations. The applicability of the SLE-diseasome was tested in different scenarios, for patient stratification analysis and for the generation and cross-cohort validation of machine learning models to predict clinical manifestations and drug response. Conclusions The SLE-diseasome offers a new SLE-specific database connecting multiple layers of database-derived biological knowledge. It is defined using a robust multicohort approach, bringing us 1 step closer to the effective use of molecular information in clinical practice through single-sample molecular scoring. References: [1.] Toro-Dominguez D. Brief Bioinform 2022;23(5):bbac332. [2.] Foroutan M. BMC Bioinformatics 2018;19:404.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,001 | 0,005 |
| Méta-épidémiologie (sens strict) | 0,002 | 0,000 |
| Méta-épidémiologie (sens large) | 0,001 | 0,002 |
| Bibliométrie | 0,008 | 0,006 |
| Études des sciences et des technologies | 0,001 | 0,000 |
| Communication savante | 0,003 | 0,002 |
| Science ouverte | 0,002 | 0,004 |
| Intégrité de la recherche | 0,001 | 0,001 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,010 | 0,005 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».