Organized Breast Screening Programs in Canada: Effect of Radiologist Reading Volumes on Outcomes
Bibliographic record
Abstract
PURPOSE: To examine retrospectively the relationship between radiologist screening program reading volumes and interpretation results. MATERIALS AND METHODS: This research project was reviewed by the University of British Columbia Research Ethics Board. Informed patient consent was not required. Data were requested from Canadian provincial screening programs for the period 1988-2000. Cancer detection rates, abnormal interpretation rates, and positive predictive values (PPVs) were calculated for individual radiologists in those programs. Multivariate Poisson mixed regression models were used to examine the effect of patient age, screening examination sequence (first or subsequent screening examination), province, radiologist reading volume, and interradiologist differences on cancer detection rate, abnormal interpretation rate, and PPV. RESULTS: The results of the interpretation of 1406678 screening mammograms by 304 radiologists from seven provincial programs were analyzed. Cancer detection rate, abnormal interpretation rate, and PPV all varied according to age of woman screened and screening sequence and across the sample of radiologists. None of the rates varied by province. Neither the cancer detection rate nor the abnormal interpretation rate varied by reading volume, but the average PPV was increased by 34% for volumes over 2000 mammograms versus volumes of 480-699 mammograms per year. There was no evidence that the magnitude of variability around the average, for radiologists reading the same volume of mammograms, varied across different volume groups for any of the outcome measures. CONCLUSION: Cancer detection did not vary with reading volume. The average PPV for individual radiologists increased as reading volume rose up to 2000 mammograms per year; it stabilized at higher volumes.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".