Abstract A052: The S-RACE Platform: Empowering Healthcare Real-World Data with Artificial Intelligence (AI) Using a Cloud-Based Approach
Bibliographic record
Abstract
Abstract Prognostic and predictive clinical decision support systems based on Real-World Evidence (RWE) data are crucial in clinical research. These systems help clinicians optimize therapeutic choices, reduce adverse events, and enhance personalized medicine. AI plays a key role in these advancements, as machine learning algorithms uncover hidden patterns beyond human capability, leading to novel clinical insights. However, the clinical translation of such algorithms depends heavily on data quality. RWE data are often unstructured, sparse, and poorly curated, requiring extensive manual processing. To address this, we developed an end-to-end data engineering pipeline to import RWE from our IT system and created machine learning models to tackle urgent clinical questions. The S-RACE Cloud-based platform, developed with Microsoft, has three main functionalities: a universal data platform (ingestion), a clinician AI hub (exploration), a data science lab (modeling), and a model registry (validation/ federated learning). The universal data platform lets investigators select patient cohorts, define data sources, and retrieve DICOM images. An on-prem anonymization engine processes data before securely transferring it. AI technologies, including Microsoft Cognitive Health Services, use natural language processing (NLP) and medical ontologies to extract relevant clinical information. Processed RWE, standardized using FHIR, are stored in a data lake and linked to specific use cases. Preliminary analyses are conducted via Microsoft Power BI, while data modeling is performed using Microsoft Azure Machine Learning Studio. Explainability techniques enhance model interpretability, and standardized templates automate documentation. Validated models will be shared internally and with the broader clinical and research communities via the clinician AI hub and using federated learning. We have integrated five major hospital IT systems into the platform. Currently, 18 clinical use cases (oncology, diabetes, multiple sclerosis, cardiovascular diseases) are under development with an overall cohort of 10k patients' data imported. At the time of writing we have developed and validated two oncological models: one for the prediction of cancer specific survival in patients with non metastatic kidney cancer at the pre-operative level and one model to predict response to (chemo)immunotherapy treatment in patients with metastatic non-small cell lung cancer. The S-RACE platform is a scalable, AI-driven approach to leveraging RWE in clinical decision-making. By integrating hospital IT systems, automating data processing, and enabling AI modeling, the platform enhances research and fosters data-driven personalized medicine. Future work will focus on expanding validated oncological AI models and facilitating their clinical adoption beyond Europe. Citation Format: Alberto Traverso, Simone Barbieri, Marco Denti, Antonio Esposito, Carlo Tacchetti. The S-RACE Platform: Empowering Healthcare Real-World Data with Artificial Intelligence (AI) Using a Cloud-Based Approach [abstract]. In: Proceedings of the AACR Special Conference in Cancer Research: Artificial Intelligence and Machine Learning; 2025 Jul 10-12; Montreal, QC, Canada. Philadelphia (PA): AACR; Clin Cancer Res 2025;31(13_Suppl):Abstract nr A052.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.009 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.002 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.005 | 0.004 |
| Open science | 0.003 | 0.004 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.021 | 0.010 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".