A73 PERFORMANCE OF ASGE AND ESGE CRITERIA FOR RISK STRATIFICATION FOR CHOLEDOCHOLITHIASIS IN A REAL-WORLD SETTING
Bibliographic record
Abstract
Abstract Background Choledocholithiasis (CDL) is a common clinical entity and can lead to serious complications, such as pancreatitis or ascending cholangitis. Endoscopic retrograde cholangio-pancreatography (ERCP) is generally the first-line procedure for definitive management of CDL. ERCP has well-established adverse events. Given the risks, patients can be stratified by likelihood of finding CDL on ERCP, thus potentially avoiding an unnecessary procedure in low probability patients. There are three commonly used criteria for this – the American Society for Gastrointestinal Endoscopy (ASGE) 2010 criteria, the ASGE 2019 criteria, and the European Society of Gastrointestinal Endoscopy (ESGE) 2019 criteria. These criteria use a mixture of biliary imaging, clinical condition, and liver biochemistry to stratify patients into low, intermediate, and high probability for CDL. Aims To test the performance characteristics of the ASGE 2010, ASGE 2019, and ESGE 2019 criteria for probability of CDL on a real-world sample. Methods We identified all adult patients who had ERCP done at our local centre for suspected CDL between 2012/01/01 and 2018/10/07. A sample of 1000 cases were chosen. We obtained the patients’ pre-procedural liver biochemistries, pre-procedural imaging in the preceding 6 months, and their ERCP reports. We used a semi-automated algorithm to determine confirmation of CDL. We inferred clinical gallstone pancreatitis using the surrogate of serum lipase at or greater than three times upper limit of normal. We could not capture clinical ascending cholangitis from the collected data. We stratified each patient according to the three guidelines and calculated their performance characteristics. Results After manually reviewing visits with incomplete ERCP or repeat ERCP, we analyzed 879 ERCP visits. There were 622 with stone or sludge found on ERCP. The performance characteristics of the high-probability and intermediate-probability criteria of the three guidelines are listed in the table below. Conclusions Our results for the 2010 ASGE guidelines high probability patients are in keeping with previous validation studies. There have been only one validation study each of the 2019 ASGE guidelines and the 2019 ESGE guidelines, and our results are different in sensitivity and negative predictive value. Future directions in refining these risk stratification tools are needed, and our project in ongoing in assessing the additional value of trends in liver biochemistry. Funding Agencies None
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.020 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.003 | 0.002 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.000 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".