Abstract A052: The S-RACE Platform: Empowering Healthcare Real-World Data with Artificial Intelligence (AI) Using a Cloud-Based Approach
Notice bibliographique
Résumé
Abstract Prognostic and predictive clinical decision support systems based on Real-World Evidence (RWE) data are crucial in clinical research. These systems help clinicians optimize therapeutic choices, reduce adverse events, and enhance personalized medicine. AI plays a key role in these advancements, as machine learning algorithms uncover hidden patterns beyond human capability, leading to novel clinical insights. However, the clinical translation of such algorithms depends heavily on data quality. RWE data are often unstructured, sparse, and poorly curated, requiring extensive manual processing. To address this, we developed an end-to-end data engineering pipeline to import RWE from our IT system and created machine learning models to tackle urgent clinical questions. The S-RACE Cloud-based platform, developed with Microsoft, has three main functionalities: a universal data platform (ingestion), a clinician AI hub (exploration), a data science lab (modeling), and a model registry (validation/ federated learning). The universal data platform lets investigators select patient cohorts, define data sources, and retrieve DICOM images. An on-prem anonymization engine processes data before securely transferring it. AI technologies, including Microsoft Cognitive Health Services, use natural language processing (NLP) and medical ontologies to extract relevant clinical information. Processed RWE, standardized using FHIR, are stored in a data lake and linked to specific use cases. Preliminary analyses are conducted via Microsoft Power BI, while data modeling is performed using Microsoft Azure Machine Learning Studio. Explainability techniques enhance model interpretability, and standardized templates automate documentation. Validated models will be shared internally and with the broader clinical and research communities via the clinician AI hub and using federated learning. We have integrated five major hospital IT systems into the platform. Currently, 18 clinical use cases (oncology, diabetes, multiple sclerosis, cardiovascular diseases) are under development with an overall cohort of 10k patients' data imported. At the time of writing we have developed and validated two oncological models: one for the prediction of cancer specific survival in patients with non metastatic kidney cancer at the pre-operative level and one model to predict response to (chemo)immunotherapy treatment in patients with metastatic non-small cell lung cancer. The S-RACE platform is a scalable, AI-driven approach to leveraging RWE in clinical decision-making. By integrating hospital IT systems, automating data processing, and enabling AI modeling, the platform enhances research and fosters data-driven personalized medicine. Future work will focus on expanding validated oncological AI models and facilitating their clinical adoption beyond Europe. Citation Format: Alberto Traverso, Simone Barbieri, Marco Denti, Antonio Esposito, Carlo Tacchetti. The S-RACE Platform: Empowering Healthcare Real-World Data with Artificial Intelligence (AI) Using a Cloud-Based Approach [abstract]. In: Proceedings of the AACR Special Conference in Cancer Research: Artificial Intelligence and Machine Learning; 2025 Jul 10-12; Montreal, QC, Canada. Philadelphia (PA): AACR; Clin Cancer Res 2025;31(13_Suppl):Abstract nr A052.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,004 | 0,009 |
| Méta-épidémiologie (sens strict) | 0,001 | 0,001 |
| Méta-épidémiologie (sens large) | 0,001 | 0,001 |
| Bibliométrie | 0,001 | 0,002 |
| Études des sciences et des technologies | 0,001 | 0,001 |
| Communication savante | 0,005 | 0,004 |
| Science ouverte | 0,003 | 0,004 |
| Intégrité de la recherche | 0,001 | 0,002 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,021 | 0,010 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».