MétaCan
Menu
Back to cohort
Record W2265124588 · doi:10.1177/0969141315600003

Study designs for determining and comparing sensitivities of disease screening tests

2015· article· en· W2265124588 on OpenAlexaff
Philip C. Prorok, Barnett S. Kramer, Anthony B. Miller

Bibliographic record

VenueJournal of Medical Screening · 2015
Typearticle
Languageen
FieldBiochemistry, Genetics and Molecular Biology
TopicMolecular Biology Techniques and Applications
Canadian institutionsUniversity of Toronto
Fundersnot available
KeywordsOverdiagnosisSensitivity (control systems)MedicineTest (biology)Randomized controlled trialCohortScreening testStatisticsInternal medicineMathematicsPediatrics

Abstract

fetched live from OpenAlex

OBJECTIVE: To investigate the capability of various study designs to determine the sensitivity of a disease screening test. METHODS: Quantities that can be calculated from these designs were derived and examined for their relationship to true sensitivity (the ability to detect unrecognized disease that would surface clinically in the absence of screening) and overdiagnosis. RESULTS: To examine the sensitivity of one test, the single cohort design, in which all participants receive the test, is particularly weak, providing only an upper bound on the true sensitivity, and yields no information about overdiagnosis. A randomized design, with one control arm and participants tested in the other, that includes sufficient post-screening follow-up, allows calculation of bounds on, and an approximation to, true sensitivity and also determination of overdiagnosis. Without follow-up, bounds on the true sensitivity can be calculated. To compare two tests, the single cohort paired design in which all participants receive both tests is precarious. The three arm randomized design with post screening follow-up is preferred, yielding an approximation to the true sensitivity, bounds on the true sensitivity, and the extent of overdiagnosis of each test. Without post screening follow-up, bounds on the true sensitivities can be calculated. When an unscreened control arm is not possible, the two-arm randomized design is recommended. Individual test sensitivities cannot be determined, but with sufficient post-screening follow-up, an order relationship can be established, as can the difference in overdiagnosis between the two tests.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.498
metaresearch head score (Gemma)0.666
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch
Consensus categoriesMetaresearch
DomainCandidate signal: Methods · Consensus signal: Methods
Study designCandidate signal: Theoretical or conceptual · Consensus signal: none
GenreCandidate signal: Methods · Consensus signal: Methods
Teacher disagreement score0.502
Threshold uncertainty score0.619

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.4980.666
Meta-epidemiology (narrow)0.0030.002
Meta-epidemiology (broad)0.0050.007
Bibliometrics0.0030.002
Science and technology studies0.0010.004
Scholarly communication0.0040.003
Open science0.0030.003
Research integrity0.0060.003
Insufficient payload (model declined to judge)0.0040.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.110
GPT teacher head0.375
Teacher spread0.265 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.

Study designTheoretical or conceptual
DomainMethods
GenreMethods

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations4
Published2015
Admission routes1
Has abstractyes

Explore more

Same venueJournal of Medical ScreeningSame topicMolecular Biology Techniques and ApplicationsFrench-language works237,207