MétaCan
Menu
Back to cohort
Record W4392966801 · doi:10.1101/2024.03.15.24303032

New implementation of data standards for AI in oncology. Experience from the EuCanImage project

2024· preprint· en· W4392966801 on OpenAlexaff
Teresa García‐Lezana, Maciej Bobowicz, Santiago Frid, Michael Rutherford, Mikel Recuero, Katrine Riklund, Aldar Cabrelles, Marlena Rygusik, Lauren A. Fromont, Roberto Francischello, Emanuele Neri, Salvador Capella-Gutiérrez, Arcadi Navarro, Fred Prior, Jonathan P. Bona, Pilar Nicolás Jiménez, Martijn P. A. Starmans, Karim Lekadir, Jordi Rambla

Bibliographic record

VenuemedRxiv · 2024
Typepreprint
Languageen
FieldMedicine
TopicRadiomics and Machine Learning in Medical Imaging
Canadian institutionsArtificial Intelligence in Medicine (Canada)
FundersAgencia Estatal de InvestigaciónMinisterio de Ciencia e InnovaciónEuskal Herriko UnibertsitateaEusko JaurlaritzaGeneralitat de CatalunyaEuropean CommissionCentres de Recerca de Catalunya
KeywordsPrecision oncologyMedical physicsComputer scienceData scienceOncologyMedicineInternal medicineCancer

Abstract

fetched live from OpenAlex

ABSTRACT Background An unprecedented amount of personal health data, with the potential to revolutionise precision medicine, is generated at healthcare institutions worldwide. The exploitation of such data using artificial intelligence relies on the ability to combine heterogeneous, multicentric, multimodal and multiparametric data, as well as thoughtful representation of knowledge and data availability. Despite these possibilities, significant methodological challenges and ethico-legal constraints still impede the real-world implementation of data models. Technical details The EuCanImage is an international consortium aimed at developing AI algorithms for precision medicine in oncology and enabling secondary use of the data based on necessary ethical approvals. The use of well-defined clinical data standards to allow interoperability was a central element within the initiative. The consortium is focused on three different cancer types and addresses seven unmet clinical needs. We have conceived and implemented an innovative process to capture clinical data from hospitals, transform it into the newly developed EuCanImage data models and then store the standardised data in permanent repositories. This new workflow combines recognized software (REDCap for data capture), data standards (FHIR for data structuring) and an existing repository (EGA for permanent data storage and sharing), with newly developed custom tools for data transformation and quality control purposes (ETL pipeline, QC scripts) to complement the gaps. Conclusion This article synthesises our experience and procedures for healthcare data interoperability, standardisation and reproducibility.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.134
metaresearch head score (Gemma)0.105
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: none
GenreCandidate signal: Methods · Consensus signal: Methods
Teacher disagreement score0.134
Threshold uncertainty score0.706

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.1340.105
Meta-epidemiology (narrow)0.0010.001
Meta-epidemiology (broad)0.0010.001
Bibliometrics0.0030.005
Science and technology studies0.0010.004
Scholarly communication0.0090.012
Open science0.0050.010
Research integrity0.0020.004
Insufficient payload (model declined to judge)0.0030.002

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.069
GPT teacher head0.483
Teacher spread0.415 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designNot applicable
Domainnot available
GenreMethods

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations3
Published2024
Admission routes1
Has abstractyes

Explore more

Same venuemedRxivSame topicRadiomics and Machine Learning in Medical ImagingFrench-language works237,207