MétaCan
Menu
← Back to cohort
Record W4221157608 · doi:10.1088/1361-6560/ac8044

OpenKBP-Opt: an international and reproducible evaluation of 76 knowledge-based planning pipelines

2022· article· en· W4221157608 on OpenAlexafffund
Aaron Babier, Rafid Mahmood, Victor Gabriel Leandro Alves, Ana María Barragán Montero, Joel Beaudry, Carlos Cárdenas, Yankui Chang, Zijie Chen, Jaehee Chun, Kelly E. Diaz, Harold David Eraso, Erik Faustmann, Sibaji Gaj, Skylar Gay, Mary Gronberg, Bingqi Guo, Gerd Heilemann, Sanchit Hira, Yuliang Huang, Fuxin Ji, Dashan Jiang, Jean Carlo Jimenez Giraldo, Hoyeon Lee, Jun Lian, Shuolin Liu, Keng‐Chi Liu, K. Miki, Kunio Nakamura, Tucker Netherton, Dan Nguyen, Hamidreza Nourzadeh, Alexander F. I. Osman, José Darío Quinto Muñoz, Christian Ramsl, Dong Joo Rhee, Juan David Rodriguez, Hongming Shan, Jeffrey V. Siebers, Mumtaz Hussain Soomro, Kay Sun, Andrés Hoyos, Carlos Valderrama, Rob Verbeek, Enpei Wang, Siri Willems, Xuanang Xu, L. Yuan, Simeng Zhu, Kevin L. Moore, Thomas G. Purdie, Andrea McNiven, Timothy C. Y. Chan

Bibliographic record

VenuePhysics in Medicine and Biology · 2022
Typearticle
Languageen
FieldPhysics and Astronomy
TopicAdvanced Radiotherapy Techniques
Canadian institutionsPrincess Margaret Cancer CentreVector InstituteUniversity of Toronto
FundersNational Cancer InstituteNatural Sciences and Engineering Research Council of Canada
KeywordsWilcoxon signed-rank testQuality assuranceVoxelComputer scienceTest planRadiation treatment planningReference dosePipeline (software)Range (aeronautics)MathematicsStatisticsMedicineArtificial intelligenceRadiation therapyOperations managementEngineeringSurgery

Abstract

fetched live from OpenAlex

Abstract Objective. To establish an open framework for developing plan optimization models for knowledge-based planning (KBP). Approach. Our framework includes radiotherapy treatment data (i.e. reference plans) for 100 patients with head-and-neck cancer who were treated with intensity-modulated radiotherapy. That data also includes high-quality dose predictions from 19 KBP models that were developed by different research groups using out-of-sample data during the OpenKBP Grand Challenge. The dose predictions were input to four fluence-based dose mimicking models to form 76 unique KBP pipelines that generated 7600 plans (76 pipelines × 100 patients). The predictions and KBP-generated plans were compared to the reference plans via: the dose score, which is the average mean absolute voxel-by-voxel difference in dose; the deviation in dose-volume histogram (DVH) points; and the frequency of clinical planning criteria satisfaction. We also performed a theoretical investigation to justify our dose mimicking models. Main results. The range in rank order correlation of the dose score between predictions and their KBP pipelines was 0.50–0.62, which indicates that the quality of the predictions was generally positively correlated with the quality of the plans. Additionally, compared to the input predictions, the KBP-generated plans performed significantly better ( P < 0.05; one-sided Wilcoxon test) on 18 of 23 DVH points. Similarly, each optimization model generated plans that satisfied a higher percentage of criteria than the reference plans, which satisfied 3.5% more criteria than the set of all dose predictions. Lastly, our theoretical investigation demonstrated that the dose mimicking models generated plans that are also optimal for an inverse planning model. Significance. This was the largest international effort to date for evaluating the combination of KBP prediction and optimization models. We found that the best performing models significantly outperformed the reference dose and dose predictions. In the interest of reproducibility, our data and code is freely available.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.020
metaresearch head score (Gemma)0.043
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch
Consensus categoriesnone
DomainCandidate signal: Evaluation · Consensus signal: none
Study designCandidate signal: Simulation or modeling · Consensus signal: Simulation or modeling
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.980
Threshold uncertainty score0.104

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0200.043
Meta-epidemiology (narrow)0.0020.001
Meta-epidemiology (broad)0.0010.002
Bibliometrics0.0030.002
Science and technology studies0.0010.001
Scholarly communication0.0020.003
Open science0.0030.005
Research integrity0.0020.002
Insufficient payload (model declined to judge)0.0030.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.287
GPT teacher head0.496
Teacher spread0.210 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

Study designSimulation or modeling
DomainEvaluation
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations1
Published2022
Admission routes2
Has abstractyes

Explore more

Same venuePhysics in Medicine and Biology→Same topicAdvanced Radiotherapy Techniques→French-language works237,207→