Exploring the Role of MRCP+ for Enhancing Detection of High-Grade Strictures in Primary Sclerosing Cholangitis
Bibliographic record
Abstract
Background: Identifying high-grade strictures (HGS) in patients with primary sclerosing cholangitis (PSC) relies upon subjective assessments of magnetic resonance cholangiopancreatography (MRCP). Quantitative MRCP (MRCP+) provides objective evaluation of MRCP examinations, which may help make these assessments more consistent and improve patient management and selection for intervention. We evaluated the impact of MRCP+ on clinicians’ confidence in diagnosing HGS in patients with PSC. Methods: Three expert abdominal radiologists independently assessed 28 patients with PSC. Radiological reads of MRCPs were performed twice, in a random order, three weeks apart, then a third time with MRCP+. HGS presence was recorded on semi-quantitative confidence scales. The cases where readers definitively agreed on presence/absence of HGS were used to assess inter- and intra-reader agreement and confidence. Results: When using MRCP alone, high intra-reader agreement was observed in identifying HGS within both intra- and extrahepatic ducts (64.3% and 70.8%, respectively), while inter-reader agreement was significantly lower for intrahepatic ducts (42.9%) than extrahepatic ducts (66.1%) (p < 0.01). Using MRCP+ in the third read significantly improved inter-reader agreement for intrahepatic HGS detection to 67.9% versus baseline reads (p = 0.02) and was comparable with extrahepatic ducts. Reader confidence tended to increase when supplemented with MRCP+, and inter-reader variability decreased. MRCP+ metrics had good performance in identifying HGS in both extra-hepatic (AUC:0.85) and intra-hepatic ducts (AUC:0.75). Conclusions: MRCP evaluation supported by quantitative metrics tended to increase individual reader confidence and reduce inter-reader variability for detecting HGS. Our results indicate that MRCP+ might help standardize MRCP assessment and subsequent management for patients with PSC.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.014 | 0.051 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".