META-ANALYSIS OF THE DIAGNOSTIC ACCURACY OF SCREENING TESTS FOR COLORECTAL CANCER
Bibliographic record
Abstract
Purpose: To conduct a meta-analysis on the diagnostic accuracy of five screening tests for colorectal cancer (CRC): fecal occult blood test (FOBT), double-contrast barium enema (DCBE), flexible sigmoidoscopy (FSIG), conventional colonoscopy (COL) and computed tomography colonoscopy (CTCOL). Methods: A literature search was carried out in MEDLINE for each test. Articles were reviewed by two independent reviewers. Inclusion criteria were: 1) RCTs or observational studies of CRC screening; 2) patients with low to average risk of CRC; 3) complete data to calculate sensitivity and specificity. Exclusion criteria were: 1) non-peer reviewed articles; 2) articles whose primary aim was not to assess CRC screening; 3) articles not in English/French; 4) articles published prior to 1975; 5) high risk screening populations. Weighted linear regression was used to identify significant covariates. Sensitivity and specificity were pooled for relevant subgroups. Results: The initial literature search found 399 articles for FOBT, 253 for DCBE, 394 for FSIG, 434 for COL, and 345 for CTCOL. Of these, 12, 8, 10, 8, and 13 articles respectively, were included in the final analysis. With the exception of colonoscopy the remaining tests showed evidence of heterogeneity and threshold effect. Significant covariates included study design and type of FOBT. Conclusions: When heterogeneity is present within test groups, results from pooled sensitivity and specificity can be misleading. A planned future step is to estimate diagnostic odds ratios and build summary ROC curves which are more reliable estimates of test accuracy for evidence synthesis.Table: Polled Sensitivity and Specificity for 5 CRC Screening Tests
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".