MétaCan
Menu
Back to cohort
Record W4409559242 · doi:10.1007/s00261-025-04941-9

Accuracy, repeatability, reproducibility and reference ranges of primary sclerosing cholangitis specific biomarkers from quantitative MRCP

2025· article· en· W4409559242 on OpenAlexaff
Mukesh G. Harisinghani, Tom R. Davis, George Ralli, Carlos Noronha Ferreira, Bruno Paun, Andrea Borghetto, Andrea Dennis, Kartik Jhaveri, Filippo Del Grande, Sarah L. Finnegan, Michele Pansini

Bibliographic record

VenueAbdominal Radiology · 2025
Typearticle
Languageen
FieldMedicine
TopicGallbladder and Bile Duct Disorders
Canadian institutionsUniversity of Toronto
Fundersnot available
KeywordsRepeatabilityReproducibilityMedicineImaging phantomNuclear medicineScannerCoefficient of variationLimits of agreementRadiologyMathematicsStatisticsComputer scienceArtificial intelligence

Abstract

fetched live from OpenAlex

PURPOSE: To assess the repeatability and reproducibility of quantitative MRCP-derived metrics generated from MRCP + software, designed for assessing biliary tree health. METHODS: Metric accuracy was assessed using a 3D-printed phantom containing 20 tubes with sinusoidally-varying diameters, simulating strictures and dilatations along ducts. Data from 80 participants (60 healthy volunteers and 20 with liver disease) was analysed in total. Repeatability and reproducibility of the quantitative metrics were assessed on Siemens, GE and Philips scanners at both 1.5T and 3T. All subjects were scanned on a Siemens Prisma 3T scanner which acted as the reference scanner. A subset of these participants also underwent scanning on the remaining scanners. Data from healthy volunteers was used to estimate the natural range of measured values (reference ranges). The reproducibility coefficient (RC) of 7 commonly reported quantitative metrics were compared between healthy controls and published values in primary sclerosing cholangitis (PSC) patients. RESULTS: The phantom analysis confirmed measurement accuracy with absolute bias of 0.0-0.1 for strictures and 0.1-0.2 for dilatations across all scanners (95% limits of agreement within ± 1.0). In vivo, RCs for the quantitative MRCP-derived metrics across the scanners ranged from: 12.4-25.4 for total number of ducts; 4.9-7.9 for number of dilatations; 3.3-6.5 for number of strictures; 4.6-9.8 mm for total length of dilatations; 26.5-51.7 mm for total length of strictures; and 4.4-6.8 for number of ducts with a stricture or dilatation. Repeatability on the same scanner was generally better than comparisons across scanners. Six metrics demonstrated sufficient cross-scanner reproducibility to distinguish healthy volunteers from PSC patients. CONCLUSION: The precision of quantitative MRCP-derived metrics were sufficient to differentiate PSC and healthy subjects and should be well suited for multi-centre trials and assessment of biliary tree health.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.001
metaresearch head score (Gemma)0.002
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.593
Threshold uncertainty score0.646

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0010.002
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0010.000
Bibliometrics0.0000.000
Science and technology studies0.0000.001
Scholarly communication0.0000.000
Open science0.0000.000
Research integrity0.0000.000
Insufficient payload (model declined to judge)0.0000.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.047
GPT teacher head0.313
Teacher spread0.266 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designObservational
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations6
Published2025
Admission routes1
Has abstractyes

Explore more

Same venueAbdominal RadiologySame topicGallbladder and Bile Duct DisordersFrench-language works237,207