Experienced Probabilities Increase Understanding of Diagnostic Test Results in Younger and Older Adults
Bibliographic record
Abstract
BACKGROUND: With advancing age, the frequency of medical screening increases. Interpreting the results of medical tests involves estimation of posterior probabilities such as positive predictive values (PPVs) and negative predictive values (NPVs). Both laypeople and experts are typically poor at estimating posterior probabilities when the relevant statistics are communicated descriptively. The current study examined whether an experience format would improve posterior probability judgments in younger and older adults, relative to a description format. METHOD: Eighty younger (ages 17-34 y) and 80 older adults (ages 65-87 y) completed an experimental task in which information about medical screening tests for 2 fictitious diseases was presented either through description or experience. Participants in the descriptive format read a passage containing statistical information, whereas participants in the experience format viewed a slideshow of representative cases that illustrated the relative frequency of the disease as well as the relative frequency of positive and negative test results. RESULTS: Both younger and older adults made more accurate posterior probability estimates in the experience format, relative to the description format. In the descriptive format, PPVs were overestimated and NPVs were underestimated. Regardless of format type, participants reported that they would prefer to rely on a physician to make medical decisions on their behalf compared with themselves. DISCUSSION: These findings are indicative of a description-experience gap in Bayesian inference, and they suggest possible avenues for enhancing medical risk communication for both younger and older patients.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.235 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".