MétaCan
Menu
← Back to cohort

Abstract B021: Current oncological large language model research lacks reproducibility, transparency, and long term support

2025· article· en· W4412163903 on OpenAlexaboutno aff
Tolou Shadbahr, Antti Rannikko, Tuomas Mirtti, Teemu D. Laajala

Bibliographic record

VenueClinical Cancer Research · 2025
Typearticle
Languageen
FieldMedicine
TopicRadiomics and Machine Learning in Medical Imaging
Canadian institutionsnot available
Fundersnot available
KeywordsTerm (time)Transparency (behavior)Current (fluid)ReproducibilityMedicineMedical physicsComputer scienceIntensive care medicineEngineeringComputer securityPhysicsStatisticsMathematics

Abstract

fetched live from OpenAlex

Abstract Large Language Models (LLMs) have been adopted increasingly in oncology, for example, in structuring data from clinical notes, inferring diagnoses from free text or imaging data, and anonymizing of data. Due to the rapid development pace of LLMs, best practices for conducting and reporting oncological research in these applications have yet to be fully established.We queried PubMed for oncology-related LLM research with the last cutoff set at Dec 31st 2024. We investigated 179 papers. Of these, 131 were removed due to omission criteria, and 48 were structured and reported here. Inclusion criteria were oncology-related research and full research articles. Structured fields included date of submission, acceptance, and publishing, the granularity of model reporting (model family, model snapshot), reporting of key LLM model parameters, availability of source code and data, and programming language and API details. We noted an almost exponential growth of LLM-related publications in oncology, with a relatively short time from authors’ submission to publicly available publication (median 3.7 months, IQR 2.5-5.9 months). Interestingly, despite the relatively short processing time, in 25% of cases, the exact model essential to the publication had been deprecated by the model service providers or a newer version was available at the time of publishing. 35.4% of published research relied solely on a graphical user-interface (GUI) of LLMs such as ChatGPT, while 37.5% reported programmatically API-use, with Python as the most common language. While most publications either fully or partially reported the utilized prompts (75%), only 22.9% reported the exact key model parameters, such as temperature. Even when the temperature parameter was available, 45.4% of these publications used a temperature value larger than 0, resulting in more stochastic answers. Source code was made publicly available in 18.7% of publications that reported using a programming language such as Python or R. While practically all publications (97.9%) reported the used model families such as GPT-4o, Claude 3.5 Sonnet or Llama 3-70B, only 27% reported the exact model snapshot usage such as GPT-4o with snapshot options available for May 13th, August 6th or November 20th in 2024. We exemplify and report shortcomings of recent LLM adoption in oncological research. To alleviate these issues, we propose a checklist to improve reproducibility, transparency, and longevity of LLM research directed at researchers and journals. We propose the following preliminary checklist: exact reporting of model snapshot and model parameter bound to a specific snapshot instead of latest release, API usage instead of GUI chatbots, temperature-parameter equal to 0, assessment of variability across runs, session restarts to avoid biases, and caution in researching models that are bound to be deprecated due to the short turn-around time in LLMs. Additionally, rigorous prompt engineering and especially few-shot learning show potential in optimizing interactions with LLMs, also in oncology. Citation Format: Tolou Shadbahr, Antti S. Rannikko, Tuomas Mirtti, Teemu D. Laajala. Current oncological large language model research lacks reproducibility, transparency, and long term support [abstract]. In: Proceedings of the AACR Special Conference in Cancer Research: Artificial Intelligence and Machine Learning; 2025 Jul 10-12; Montreal, QC, Canada. Philadelphia (PA): AACR; Clin Cancer Res 2025;31(13_Suppl):Abstract nr B021.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.188
metaresearch head score (Gemma)0.577
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch
Consensus categoriesnone
DomainCandidate signal: Reproducibility · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: none
Teacher disagreement score0.812
Threshold uncertainty score0.996

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.1880.577
Meta-epidemiology (narrow)0.0010.002
Meta-epidemiology (broad)0.0020.005
Bibliometrics0.0080.015
Science and technology studies0.0020.004
Scholarly communication0.0200.020
Open science0.0050.008
Research integrity0.0030.005
Insufficient payload (model declined to judge)0.0420.018

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.324
GPT teacher head0.634
Teacher spread0.311 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

Study designObservational
DomainReproducibility
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2025
Admission routes1
Has abstractyes

Explore more

Same venueClinical Cancer Research→Same topicRadiomics and Machine Learning in Medical Imaging→French-language works237,207→