Creating Robust Predictive Radiomic Models for Data From Independent Institutions Using Normalization
Bibliographic record
Abstract
Purpose: The distribution of a radiomic feature can differ between two institutions due to, for example, different image acquisition parameters, imaging systems, and contouring (i.e., tumor delineation) variations between clinicians. We aimed to develop effective statistical methods to successfully apply a radiomics-based predictive model to an external dataset. Theory: Two common feature normalization methods, rescaling and standardization, were evaluated for suitability in reducing feature variability between institutions. Standardization was chosen as the preferred approach, since rescaling was more sensitive to statistical outliers, and potentially reduced the discrimination power of a feature. It was also demonstrated why a dataset needs to be balanced between positive and negative outcomes before standardization is applied to it. Methods: In this paper, the novelty and power of the developed method for improved application of radiomics models on external datasets is tied to finding the normalization transformations separately for each independent set. The clinical effectiveness of the normalization method was shown using magnetic resonance images of primary uterine adenocarcinoma. Feature selection was done using 94 samples (Institution X), and feature testing was done using 63 samples (Institution Y). The outcomes studied were lymphovascular space invasion and cancer staging. Logistic regression was used to obtain the prediction accuracy of a feature. Promising radiomic features were defined as those with AUC > 0.75 in the training set. Results: When comparing the prediction accuracy, F-score, and Matthews correlation coefficient (MCC) of promising radiomic features in the testing set with and without standardization, there was an improvement due to standardization. For cancer stage prediction, average accuracy for all promising features rose from 0.64 to 0.72, average F-score from 0.48 to 0.71, and average MCC from 0.34 to 0.44 (p-5). Furthermore, when applying standardization, the ratio of sensitivity to specificity was close to unity in the testing set, comparable to the ratio in the training set. Without standardization, this ratio deviated significantly from unity in the testing set. Conclusions: Applying feature standardization separately for each independent set using imbalance adjustments was shown to improve the predictive ability of radiomic models when applied to a dataset from an external institution.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.017 | 0.037 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.003 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.002 | 0.002 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".