MétaCan
Menu
Back to cohort
Record W4410344814 · doi:10.1093/gigascience/giae101

New implementation of data standards for AI in oncology: Experience from the EuCanImage project

2024· article· en· W4410344814 on OpenAlexaff
Teresa García‐Lezana, Maciej Bobowicz, Santiago Frid, Michael Rutherford, Mikel Recuero, Katrine Riklund, Aldar Cabrelles, Marlena Rygusik, Lauren A. Fromont, Roberto Francischello, Emanuele Neri, Salvador Capella-Gutiérrez, Arcadi Navarro, Fred Prior, Jonathan P. Bona, Pilar Nicolás Jiménez, Martijn P. A. Starmans, Karim Lekadir, Jordi Rambla

Bibliographic record

VenueGigaScience · 2024
Typearticle
Languageen
FieldMedicine
TopicArtificial Intelligence in Healthcare and Education
Canadian institutionsArtificial Intelligence in Medicine (Canada)
FundersAgencia Estatal de InvestigaciónEuskal Herriko UnibertsitateaHorizon 2020 Framework ProgrammeEusko JaurlaritzaEuropean CommissionMinisterio de Ciencia e InnovaciónCentres de Recerca de Catalunya
KeywordsWorkflowStandardizationInteroperabilityElectronic data captureComputer scienceData scienceData qualityData sharingAutomatic identification and data captureData dictionaryHealth carePrecision medicineMetadataMedicineDatabaseWorld Wide WebClinical trialEngineering

Abstract

fetched live from OpenAlex

BACKGROUND: An unprecedented amount of personal health data, with the potential to revolutionize precision medicine, is generated at health care institutions worldwide. The exploitation of such data using artificial intelligence (AI) relies on the ability to combine heterogeneous, multicentric, multimodal, and multiparametric data, as well as thoughtful representation of knowledge and data availability. Despite these possibilities, significant methodological challenges and ethicolegal constraints still impede the real-world implementation of data models. TECHNICAL DETAILS: The EuCanImage is an international consortium aimed at developing AI algorithms for precision medicine in oncology and enabling secondary use of the data based on necessary ethical approvals. The use of well-defined clinical data standards to allow interoperability was a central element within the initiative. The consortium is focused on 3 different cancer types and addresses 7 unmet clinical needs. We have conceived and implemented an innovative process to capture clinical data from hospitals, transform it into the newly developed EuCanImage data models, and then store the standardized data in permanent repositories. This new workflow combines recognized software (REDCap for data capture), data standards (FHIR for data structuring), and an existing repository (EGA for permanent data storage and sharing), with newly developed custom tools for data transformation and quality control purposes (ETL pipeline, QC scripts) to complement the gaps. CONCLUSION: This article synthesizes our experience and procedures for health care data interoperability, standardization, and reproducibility.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.001
metaresearch head score (Gemma)0.000
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Other design · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.909
Threshold uncertainty score0.987

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0010.000
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0000.000
Bibliometrics0.0000.000
Science and technology studies0.0000.000
Scholarly communication0.0000.000
Open science0.0000.000
Research integrity0.0000.000
Insufficient payload (model declined to judge)0.0000.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.413
GPT teacher head0.627
Teacher spread0.215 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designOther design
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations3
Published2024
Admission routes1
Has abstractyes

Explore more

Same venueGigaScienceSame topicArtificial Intelligence in Healthcare and EducationFrench-language works237,207