MétaCan
Menu
Back to cohort

Abstract A052: The S-RACE Platform: Empowering Healthcare Real-World Data with Artificial Intelligence (AI) Using a Cloud-Based Approach

2025· article· en· W4412163733 on OpenAlexaboutno aff
Alberto Traverso, Simone Barbieri, Marco Denti, Antonio Esposito, Carlo Tacchetti

Bibliographic record

VenueClinical Cancer Research · 2025
Typearticle
Languageen
FieldHealth Professions
TopicArtificial Intelligence in Healthcare
Canadian institutionsnot available
Fundersnot available
KeywordsRace (biology)Cloud computingComputer scienceHealth careArtificial intelligenceData scienceMedicineBiologyOperating systemPolitical science

Abstract

fetched live from OpenAlex

Abstract Prognostic and predictive clinical decision support systems based on Real-World Evidence (RWE) data are crucial in clinical research. These systems help clinicians optimize therapeutic choices, reduce adverse events, and enhance personalized medicine. AI plays a key role in these advancements, as machine learning algorithms uncover hidden patterns beyond human capability, leading to novel clinical insights. However, the clinical translation of such algorithms depends heavily on data quality. RWE data are often unstructured, sparse, and poorly curated, requiring extensive manual processing. To address this, we developed an end-to-end data engineering pipeline to import RWE from our IT system and created machine learning models to tackle urgent clinical questions. The S-RACE Cloud-based platform, developed with Microsoft, has three main functionalities: a universal data platform (ingestion), a clinician AI hub (exploration), a data science lab (modeling), and a model registry (validation/ federated learning). The universal data platform lets investigators select patient cohorts, define data sources, and retrieve DICOM images. An on-prem anonymization engine processes data before securely transferring it. AI technologies, including Microsoft Cognitive Health Services, use natural language processing (NLP) and medical ontologies to extract relevant clinical information. Processed RWE, standardized using FHIR, are stored in a data lake and linked to specific use cases. Preliminary analyses are conducted via Microsoft Power BI, while data modeling is performed using Microsoft Azure Machine Learning Studio. Explainability techniques enhance model interpretability, and standardized templates automate documentation. Validated models will be shared internally and with the broader clinical and research communities via the clinician AI hub and using federated learning. We have integrated five major hospital IT systems into the platform. Currently, 18 clinical use cases (oncology, diabetes, multiple sclerosis, cardiovascular diseases) are under development with an overall cohort of 10k patients' data imported. At the time of writing we have developed and validated two oncological models: one for the prediction of cancer specific survival in patients with non metastatic kidney cancer at the pre-operative level and one model to predict response to (chemo)immunotherapy treatment in patients with metastatic non-small cell lung cancer. The S-RACE platform is a scalable, AI-driven approach to leveraging RWE in clinical decision-making. By integrating hospital IT systems, automating data processing, and enabling AI modeling, the platform enhances research and fosters data-driven personalized medicine. Future work will focus on expanding validated oncological AI models and facilitating their clinical adoption beyond Europe. Citation Format: Alberto Traverso, Simone Barbieri, Marco Denti, Antonio Esposito, Carlo Tacchetti. The S-RACE Platform: Empowering Healthcare Real-World Data with Artificial Intelligence (AI) Using a Cloud-Based Approach [abstract]. In: Proceedings of the AACR Special Conference in Cancer Research: Artificial Intelligence and Machine Learning; 2025 Jul 10-12; Montreal, QC, Canada. Philadelphia (PA): AACR; Clin Cancer Res 2025;31(13_Suppl):Abstract nr A052.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.004
metaresearch head score (Gemma)0.009
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: Not applicable
GenreCandidate signal: Other · Consensus signal: none
Teacher disagreement score0.021
Threshold uncertainty score0.070

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0040.009
Meta-epidemiology (narrow)0.0010.001
Meta-epidemiology (broad)0.0010.001
Bibliometrics0.0010.002
Science and technology studies0.0010.001
Scholarly communication0.0050.004
Open science0.0030.004
Research integrity0.0010.002
Insufficient payload (model declined to judge)0.0210.010

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.805
GPT teacher head0.707
Teacher spread0.097 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designNot applicable
Domainnot available
GenreOther

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2025
Admission routes1
Has abstractyes

Explore more

Same venueClinical Cancer ResearchSame topicArtificial Intelligence in HealthcareFrench-language works237,207