RADCURE: An open‐source head and neck cancer CT dataset for clinical radiation therapy insights
Bibliographic record
Abstract
PURPOSE: This manuscript presents RADCURE, one of the most extensive head and neck cancer (HNC) imaging datasets accessible to the public. Initially collected for clinical radiation therapy (RT) treatment planning, this dataset has been retrospectively reconstructed for use in imaging research. ACQUISITION AND VALIDATION METHODS: RADCURE encompasses data from 3346 patients, featuring computed tomography (CT) RT simulation images with corresponding target and organ-at-risk contours. These CT scans were collected using systems from three different manufacturers. Standard clinical imaging protocols were followed, and contours were manually generated and reviewed at weekly RT quality assurance rounds. RADCURE imaging and structure set data was extracted from our institution's radiation treatment planning and oncology information systems using a custom-built data mining and processing system. Furthermore, images were linked to our clinical anthology of outcomes data for each patient and includes demographic, clinical and treatment information based on the 7th edition TNM staging system (Tumor-Node-Metastasis Classification System of Malignant Tumors). The median patient age is 63, with the final dataset including 80% males. Half of the cohort is diagnosed with oropharyngeal cancer, while laryngeal, nasopharyngeal, and hypopharyngeal cancers account for 25%, 12%, and 5% of cases, respectively. The median duration of follow-up is five years, with 60% of the cohort surviving until the last follow-up point. DATA FORMAT AND USAGE NOTES: The dataset provides images and contours in DICOM CT and RT-STRUCT formats, respectively. We have standardized the nomenclature for individual contours-such as the gross primary tumor, gross nodal volumes, and 19 organs-at-risk-to enhance the RT-STRUCT files' utility. Accompanying demographic, clinical, and treatment data are supplied in a comma-separated values (CSV) file format. This comprehensive dataset is publicly accessible via The Cancer Imaging Archive. POTENTIAL APPLICATIONS: RADCURE's amalgamation of imaging, clinical, demographic, and treatment data renders it an invaluable resource for a broad spectrum of radiomics image analysis research endeavors. Researchers can utilize this dataset to advance routine clinical procedures using machine learning or artificial intelligence, to identify new non-invasive biomarkers, or to forge prognostic models.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".