Bibliographic record
Abstract
Currently in the United States, digital mammography has almost completely replaced film-screen mammography, although it was recognized early that specificity was reduced (the numbers of normal results deemed falsely positive to the test increased) even though the sensitivity of the test was increased (the number of women found positive to the test of those who truly had the disease) (1). However, increased sensitivity in detecting disease is not necessarily accompanied by the benefit sought (ie, reduced numbers of deaths from breast cancer in screened women). Using five mathematical models from the CISNET consortium, Stout et al. in this issue of the Journal address this important issue (2). All of the models used identical data as input. Although the approaches used to model the natural history of breast cancer were different in the four models that attempted this, the fifth model begins at cancer detection and does not explicitly capture natural history. Medians were derived from the results of the different models in making the final estimates. The fact that the conclusions were derived from the application of five models makes them far more robust than if only a single model were used. This is one of the major strengths of the CISNET consortium. However, Supplementary Table 3 in Stout et al. (2) indicates that there was substantial variation in the estimated effect of screening, ranging, for example, from 23% to 56% breast cancer mortality reduction for annual digital mammography screening of women aged 40 to 74 years. It is unfortunate that the authors did not explain why the variation occurred, other than postulating in the Discussion that it may be because of different modeled effects of treatment. If that is so, perhaps the treatment parameters in the models that compute most benefit from screening need to be adjusted because evidence is accruing that advances in treatment have resulted in a negligible effect of screening on breast cancer mortality (3,4).
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.005 | 0.035 |
| Meta-epidemiology (narrow) | 0.003 | 0.001 |
| Meta-epidemiology (broad) | 0.003 | 0.002 |
| Bibliometrics | 0.005 | 0.002 |
| Science and technology studies | 0.002 | 0.002 |
| Scholarly communication | 0.008 | 0.003 |
| Open science | 0.003 | 0.002 |
| Research integrity | 0.013 | 0.016 |
| Insufficient payload (model declined to judge) | 0.034 | 0.020 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".