Application of decision rules on diagnosis and prognosis of renal colic: a systematic review and meta-analysis
Bibliographic record
Abstract
Renal colic is a prevalent emergency department presentation resulting from urolithiasis. Clinical decision rules for the diagnosis of urolithiasis were developed to help clinicians with better judgment. In this systematic review, we assessed the performance of prediction rules on urolithiasis diagnosis and prognosis. MEDLINE, Embase, Web of Science, and Scopus were searched for studies on the performance of a clinical decision tool for diagnosis or prognosis of urolithiasis. Performance and accuracy of the rules were the key outcomes of interest. Databases were searched from inception to March 2019. Of the 4980 articles reviewed, 28 studies were included in the present analysis. Twenty-one studies were on urolithiasis diagnosis (including eight studies on STONE rule), and 10 studies reported urolithiasis outcomes. Studies were at low to moderate risk of bias. The pooling of data on STONE showed that the prevalence of urolithiasis in low, moderate, and high risk groups were: 12% (95% confidence interval 9%-15%), 53% (95% confidence interval 43%-62%), and 83% (95% confidence interval 75%-91%), respectively. In the high risk score group, prevalence of clinically important alternative diagnosis was 1% (95% confidence interval 0%-2%) and 11% (95% confidence interval 8%-13%) of patients needed intervention. STONE scoring system is useful in estimating the prevalence of urolithiasis but high heterogeneity among the studies makes it unsuitable for application. Other decision tools were poorly studied and cannot be recommended for clinical use.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.010 | 0.002 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".