Mammographic density assessed on paired raw and processed digital images and on paired screen-film and digital images across three mammography systems
Bibliographic record
Abstract
BACKGROUND: Inter-women and intra-women comparisons of mammographic density (MD) are needed in research, clinical and screening applications; however, MD measurements are influenced by mammography modality (screen film/digital) and digital image format (raw/processed). We aimed to examine differences in MD assessed on these image types. METHODS: We obtained 1294 pairs of images saved in both raw and processed formats from Hologic and General Electric (GE) direct digital systems and a Fuji computed radiography (CR) system, and 128 screen-film and processed CR-digital pairs from consecutive screening rounds. Four readers performed Cumulus-based MD measurements (n = 3441), with each image pair read by the same reader. Multi-level models of square-root percent MD were fitted, with a random intercept for woman, to estimate processed-raw MD differences. RESULTS: respectively, mean √dense area difference 0.44 cm (95% CI: 0.36, 0.52)). This difference in √dense area was significant for direct digital systems (Hologic 0.50 cm (95% CI: 0.39, 0.61), GE 0.56 cm (95% CI: 0.42, 0.69)) but not for Fuji CR (0.06 cm (95% CI: -0.10, 0.23)). Additionally, within each system, reader-specific differences varied in magnitude and direction (p < 0.001). Conversion equations revealed differences converged to zero with increasing dense area. MD differences between screen-film and processed digital on the subsequent screening round were consistent with expected time-related MD declines. CONCLUSIONS: MD was slightly higher when measured on processed than on raw direct digital mammograms. Comparisons of MD on these image formats should ideally control for this non-constant and reader-specific difference.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.002 |
| Scholarly communication | 0.002 | 0.002 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".