Accuracy, repeatability, reproducibility and reference ranges of primary sclerosing cholangitis specific biomarkers from quantitative MRCP
Bibliographic record
Abstract
PURPOSE: To assess the repeatability and reproducibility of quantitative MRCP-derived metrics generated from MRCP + software, designed for assessing biliary tree health. METHODS: Metric accuracy was assessed using a 3D-printed phantom containing 20 tubes with sinusoidally-varying diameters, simulating strictures and dilatations along ducts. Data from 80 participants (60 healthy volunteers and 20 with liver disease) was analysed in total. Repeatability and reproducibility of the quantitative metrics were assessed on Siemens, GE and Philips scanners at both 1.5T and 3T. All subjects were scanned on a Siemens Prisma 3T scanner which acted as the reference scanner. A subset of these participants also underwent scanning on the remaining scanners. Data from healthy volunteers was used to estimate the natural range of measured values (reference ranges). The reproducibility coefficient (RC) of 7 commonly reported quantitative metrics were compared between healthy controls and published values in primary sclerosing cholangitis (PSC) patients. RESULTS: The phantom analysis confirmed measurement accuracy with absolute bias of 0.0-0.1 for strictures and 0.1-0.2 for dilatations across all scanners (95% limits of agreement within ± 1.0). In vivo, RCs for the quantitative MRCP-derived metrics across the scanners ranged from: 12.4-25.4 for total number of ducts; 4.9-7.9 for number of dilatations; 3.3-6.5 for number of strictures; 4.6-9.8 mm for total length of dilatations; 26.5-51.7 mm for total length of strictures; and 4.4-6.8 for number of ducts with a stricture or dilatation. Repeatability on the same scanner was generally better than comparisons across scanners. Six metrics demonstrated sufficient cross-scanner reproducibility to distinguish healthy volunteers from PSC patients. CONCLUSION: The precision of quantitative MRCP-derived metrics were sufficient to differentiate PSC and healthy subjects and should be well suited for multi-centre trials and assessment of biliary tree health.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".