How much net gain does a diagnostic imaging test provide?
Bibliographic record
Abstract
Heston 1 describes the importance of prevalence in order to judge the utility of a diagnostic test. Heston 1 recommended standardizing predictive values to a prevalence of 50%. We take this one step further and present the concept of gain of a diagnostic test and the Predictive Summary Index (PSI), which uses the prevalence of a disease in the population of interest in order to standardize the predictive values. The PSI derivation is explained, and we show that it represents the net gain in information from a diagnostic test. Heston 1 comments on an important issue when assessing diagnostics tests: the need to know the overall prevalence of the disease in the population under investigation in order to make judgment of the utility of a diagnostic test. As correctly illustrated by Heston 1, altering the pretest probability affects the posttest probability. Heston 1 suggests that a better way to present the predictive value would be to standardize it to a 50% disease prevalence in order to reduce prevalence bias when comparing diagnostic tests. We suggest taking this one step further. Why limit standardization to 50%? The new measure, the PSI 2, can be used as a relevant summary measure to show the net gain in information beyond the prevalence of the disease. The positive predictive value (PPV) could be meaningful and supply information only if it is greater than the prevalence of the disease, which could be estimated based on pretest likelihood of disease. Thus, the additional information that is obtained from a given diagnostic test is the gain in information after a positive test which is G+ = PPV − prevalence. The negative predictive value (NPV) could be meaningful and supply information only if it is greater than the prevalence of the lack of a disease (which equals 1 − the prevalence of a disease) which could be our first guess of a lack of a disease in a patient without performing any test. The gain in information after a negative test is G− = NPV− (1−prevalence). The PSI 2 takes both of these gain measures into account. PSI is calculated as: As shown, PSI is actually measuring the net gain in information (for positive and negative results) obtained for a diagnostic test beyond the prevalence of the disease. To take it one step further, in the clinical setting a False Positive Rate (FPR) of a test is FPR = 1−PPV and False Negative Rate (FNR) is FNR = 1−NPV. Thus, PSI can also be expressed as PSI = 1− (FPR+FNR). PSI is therefore a summary of the error-free diagnostic capabilities of a test for a disease or its absence. A PSI of 1 indicates an ideal test without errors in diagnosing a disease or its absence; a PSI of 0 indicates an uninformative test with an error rate that equals the diagnosis rate. PSI of −1 indicates a misleading test that always diagnoses a disease incorrectly. Based on data by Pilz et al. 3 and the example by Heston 1, we can use the data below, using cardiac magnetic resonance imaging (CMR) compared to the gold standard coronary angiography (CA) (Table 1). The sensitivity, specificity, PPV, and NPV are 84%, 55%, 20%, 96%, respectively. These predictive values are only relevant to the patient population with the overall prevalence of 38/316 = 12%. The PSI = 0.16 indicating an overall gain in information of 16%. For purposes of illustration, let's apply this to a population with prevalence of 50% using the same sensitivity and specificity. This would yield a PPV, NPV, PSI of 65%, 77%, and 0.42, respectively (thus a net gain in information of 42%). For a population with prevalence of 75% (as done in Henson 1), again using the same sensitivity and specificity, the PPV, NPV, and PSI will be 85%, 47%, and 0.32 (thus a net gain of information of 32%). As seen by this mathematical exercise, the overall gain in information from a diagnostic test is strictly dependent on prevalence. Furthermore, we suggest that the PSI is much more informative than reporting a generic standardized predictive value to an arbitrary prevalence of 50% 1, given the PSI varies based on the prevalence in the population of interest. Gilat Grunau, PhD1 Peter Grunau, MD2 Shai Linn, MD, PhD3 Jonathon Leipsic, MD1,4 1Department of Radiology University of British Columbia Vancouver, BC, Canada 2Department of Orthopedic Surgery University of British Columbia Vancouver, BC, Canada 3School of Public Health University of Haifa Haifa, Israel 4Department of Medical Imaging St Paul's Hospital Vancouver, BC, Canada
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.007 | 0.437 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.003 | 0.001 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.002 | 0.000 |
| Research integrity | 0.000 | 0.005 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".