Accuracy of Colorectal Polyp Self-Reports: Findings from the Colon Cancer Family Registry
Bibliographic record
Abstract
INTRODUCTION: Colorectal adenomas and other types of polyps are commonly used as end points or risk factors in epidemiologic studies. However, it is not known how accurately patients are able to self-report the presence or absence of adenomas following colonoscopy. METHODS: Participants in the Colon Cancer Family Registry provided self-reports of recent colorectal cancer (CRC) screening activity, and whether or not they had ever been told they had a polyp. Positive and negative predictive values for polyp self-report were calculated by comparing medical records with self-reports from 488 participants. RESULTS: The positive predictive value for self-reported polyp was 80.9%, and the negative predictive value was 85.8%. The predictive values did not differ by age group or sex, but participants with a previous diagnosis of CRC had a lower negative predictive value (76.2%) than participants with no personal history of CRC (89.0%; P = 0.04). CONCLUSIONS: Predictive values for self-reports of polyps are fairly high, but researchers needing accurate polyp data should obtain medical record confirmation. Pursuing medical records on only those participants self-reporting a polyp could result in an underestimation of the polyp prevalence in a study population.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".