Prospective Study of Biliary Strictures to Determine the Predictors of Malignancy
Bibliographic record
Abstract
BACKGROUND: There have been few prospective studies regarding the investigation of biliary strictures, principally because of rapid technological change. The present study was designed to determine the sensitivity of various imaging studies for the detection of biliary strictures. Serum biochemistry and imaging studies were evaluated for their role in distinguishing benign from malignant strictures. METHODS: Thirty-one patients with suspected noncalculus biliary obstruction were enrolled consecutively in the study. A complete biochemical profile, ultrasound, Disida scan and cholangiogram (endoscopic retrograde cholangiopancreatography [ERCP] or percutaneous cholangiogram) were obtained at study entry. Stricture etiology was determined based on cytology, biopsy and/or clinical follow-up at one year. RESULTS: Twenty-nine of 31 patients had biliary strictures, of which 15 were malignant. The mean age of the malignant cohort was 73.9 years versus 53.9 years in the benign cohort (P<0.001). Statistically significant differences between the malignant and benign groups, respectively, were as follows: alanine transaminase 235.2 versus 66.9 U/L (P=0.004), aspartate transaminase 189.8 versus 84.5 U/L (P=0.011), alkaline phosphatase 840.2 versus 361.1 U/L (P=0.002), bilirubin 317.8 versus 22.1 micromol/L (P<0. 001) and bile acids 242.5 versus 73.2 micromol/L (P=0.001). Threshold analysis using receiver operative characteristic (ROC) curves demonstrated that a bilirubin level of 75 micromol/L was most predictive of malignant strictures. Intrahepatic duct dilation was present in 93% of malignant strictures versus 36% of benign strictures (P=0.002). Common hepatic duct dilation was less discriminatory (malignant 13.5 versus benign 9.6 mm; P=0.11). Ultrasound was highly sensitive (93%) in the detection of the primary tumour in the bile duct or pancreas, or in the visualization of nodal or liver metastases. In benign disease, ultrasound failed to detect evidence of intrahepatic or extrahepatic biliary dilation in most cases. Disida scans were not able to distinguish between malignant or benign strictures and could not accurately localize the level of obstruction. The sensitivity of Disida scan for the diagnosis of obstruction was 50%. Cholangiographic characterization of strictures revealed an equal distribution of smooth (eight of 13) and irregular (five of 13) strictures in the malignant group. Ten of 13 benign strictures were characterized as smooth. Malignant strictures were significantly longer than benign ones - 30.3 versus 9.2 mm (P=0.001). Threshold analysis using ROC curves showed that strictures greater than or equal to 14 mm were predictive of malignancy (sensitivity 78%, specificity 75%, log odds ratio 11.23). CONCLUSIONS: A serum bilirubin level of 75 micromol/L or higher, or a stricture length of greater than 14 mm was highly predictive of malignancy in patients with a biliary stricture. Ultrasound was useful in predicting malignant strictures by detecting either intrahepatic duct dilation or by visualizing the tumour (primary or metastases). Strictures with a 'benign' cholangiographic appearance are frequently malignant. Disida scan did not add additional information. ERCP is necessary to diagnose benign strictures, which tend to be less extensive at presentation.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".