Study designs for determining and comparing sensitivities of disease screening tests
Bibliographic record
Abstract
OBJECTIVE: To investigate the capability of various study designs to determine the sensitivity of a disease screening test. METHODS: Quantities that can be calculated from these designs were derived and examined for their relationship to true sensitivity (the ability to detect unrecognized disease that would surface clinically in the absence of screening) and overdiagnosis. RESULTS: To examine the sensitivity of one test, the single cohort design, in which all participants receive the test, is particularly weak, providing only an upper bound on the true sensitivity, and yields no information about overdiagnosis. A randomized design, with one control arm and participants tested in the other, that includes sufficient post-screening follow-up, allows calculation of bounds on, and an approximation to, true sensitivity and also determination of overdiagnosis. Without follow-up, bounds on the true sensitivity can be calculated. To compare two tests, the single cohort paired design in which all participants receive both tests is precarious. The three arm randomized design with post screening follow-up is preferred, yielding an approximation to the true sensitivity, bounds on the true sensitivity, and the extent of overdiagnosis of each test. Without post screening follow-up, bounds on the true sensitivities can be calculated. When an unscreened control arm is not possible, the two-arm randomized design is recommended. Individual test sensitivities cannot be determined, but with sufficient post-screening follow-up, an order relationship can be established, as can the difference in overdiagnosis between the two tests.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.498 | 0.666 |
| Meta-epidemiology (narrow) | 0.003 | 0.002 |
| Meta-epidemiology (broad) | 0.005 | 0.007 |
| Bibliometrics | 0.003 | 0.002 |
| Science and technology studies | 0.001 | 0.004 |
| Scholarly communication | 0.004 | 0.003 |
| Open science | 0.003 | 0.003 |
| Research integrity | 0.006 | 0.003 |
| Insufficient payload (model declined to judge) | 0.004 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".