Diagnostic utility of alarm features for colorectal cancer: systematic review and meta-analysis
Bibliographic record
Abstract
OBJECTIVE: Colorectal cancer is the second most common cause of cancer death in Europe and North America. Alarm features are used to prioritize access to urgent investigation, but there is little information concerning their utility in the diagnosis of colorectal cancer. METHODS: A systematic review and meta-analysis of the published literature was carried out to assess the diagnostic accuracy of alarm features in predicting colorectal cancer. Primary or secondary care-based studies in unselected cohorts of adult patients with lower gastrointestinal symptoms were identified by searching MEDLINE, EMBASE and CINAHL (up to October 2007). The main outcome measures were accuracy of alarm features or statistical models in predicting the presence of colorectal cancer after investigation. Data were pooled to estimate sensitivity, specificity, and positive and negative likelihood ratios. The quality of the included studies was assessed according to predefined criteria. RESULTS: Of 11 169 studies identified, 205 were retrieved for evaluation. Fifteen studies were eligible for inclusion, evaluating 19 443 patients, with a pooled prevalence of colorectal carcinoma of 6% (95% CI 5% to 8%). Pooled sensitivity of alarm features was poor (5% to 64%) but specificity was >95% for dark red rectal bleeding and abdominal mass, suggesting that the presence of either rules the diagnosis of colorectal cancer in. Statistical models had a sensitivity of 90%, but poor specificity. CONCLUSIONS: Most alarm features had poor sensitivity and specificity for the diagnosis of colorectal carcinoma, whilst statistical models performed better in terms of sensitivity. Future studies should examine the utility of dark red rectal bleeding and abdominal mass, and concentrate on maximising specificity when validating statistical models.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.022 | 0.081 |
| Meta-epidemiology (narrow) | 0.003 | 0.002 |
| Meta-epidemiology (broad) | 0.018 | 0.036 |
| Bibliometrics | 0.007 | 0.007 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.003 | 0.002 |
| Open science | 0.003 | 0.001 |
| Research integrity | 0.003 | 0.002 |
| Insufficient payload (model declined to judge) | 0.003 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".