Using Medical Imaging Effective Dose in Deep Learning Models: Estimation and Evaluation
Bibliographic record
Abstract
Accurately estimating patient exposure is a fundamental concern when modeling the radiation risk from medical imaging. Inaccurate estimation techniques produce misleading instances that have a cascading effect on the performance of risk models. A commonly used method of estimating exposure is using mean effective dose (ED) values from the literature. However, the predictive power of literature values to impute patients ED has not been investigated. In this article, we present a comparative analysis between two methods to estimate ED for computed tomography (CT) and X-ray (XR) scans: 1) mean EDs from the literature and 2) calculated dose estimates from imaging scans using ED tools. We also used adversarial machine learning to demonstrate how the difference between estimation methods impacts a proof-of-concept deep learning model. The study cohort had 39 909 medical imaging scans (7427 CT and 32 482 XR scans) from a stratified random sample of 2000 patients from four hospitals in Hamilton, Canada over ten years. Our results showed moderate increases in the mean ED compared to the literature across all exam types. However, using mean values reported in the literature underestimated patients' total ED over the study period. Our results also demonstrated that the differences between the estimation methods were enough to cause model misclassifications. The results of our study demonstrate the challenges of using mean ED values from the literature to estimate patient medical imaging exposure. There is a need to develop novel imputation methods to estimate patients' EDs from medical imaging.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".