Optimization of Gross Tumor Volume Dilation for Improved Radiomic Prediction of HPV Status in Head and Neck Cancer
Bibliographic record
Abstract
Accuracy of head and neck cancer (HNC) treatment and prognostication using human papillomavirus (HPV) status is of great significance. We examine if dilation of the gross tumor volume (GTV) for peritumoral inclusion improves HPV prediction using computed tomography (CT)-based radiomic features. 1726 subjects from three publicly available datasets were evaluated. The GTVs were dilated systematically from 1 to 34 mm. Radiomic features were extracted from original and dilated structures after resampling and intensity clipping. Feature selection was performed using Boruta, Minimum Redundancy Maximum Relevance (MRMR), and Recursive Feature Elimination (RFE) methods and classification was performed using Support Vector Machine (SVM) and Gradient-Boosted (GB) machine learning models under a 5 -fold nested crossvalidation protocol. AUC, sensitivity and specificity were evaluated. The results showed that the area under the curve (AUC) improved from 0.73 at 0 mm dilation to over 0.82 at 20 mm, after which it plateaued. Balanced accuracy (BAC) showed a similar trend, increasing from approximately 0.68 to 0.74 at 20 mm. Sensitivity and specificity also improved from 0.65 at baseline to 0.75 and 0.80, respectively, at 20 mm, with no further gains observed up to 34 mm. Models using SVM with Boruta or MRMR demonstrated more consistent performance compared to those with RFE, which showed greater variability beyond 20 mm dilation, highlighting the benefit of incorporating peritumoral tissue up to an optimal margin. These findings suggest that optimized GTV dilation enhances non-invasive HPV prediction and may serve as a valuable adjunct to histopathological testing in HNC patients.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.004 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".