MétaCan
Menu
Back to cohort
Record W3118767406 · doi:10.1093/ecco-jcc/jjaa267

A Reliability Study: Strong Inter-Observer Agreement of an Expert Panel for Intestinal Ultrasound in Ulcerative Colitis

2020· article· en· W3118767406 on OpenAlexaff
F de Voogd, Rune Wilkens, K Gecse, Mariangela Allocca, Kerri L. Novak, Cathy Lu, Geert D’Haens, Christian Maaser

Bibliographic record

VenueJournal of Crohn s and Colitis · 2020
Typearticle
Languageen
FieldBiochemistry, Genetics and Molecular Biology
TopicInflammatory Bowel Disease
Canadian institutionsUniversity of Calgary
Fundersnot available
KeywordsMedicineUlcerative colitisUltrasoundKappaGrading (engineering)GastroenterologyRadiologyInflammatory bowel diseaseInternal medicineDisease

Abstract

fetched live from OpenAlex

BACKGROUND: Intestinal ultrasound [IUS] is a promising and non-invasive cross-sectional imaging modality in the diagnosis and monitoring of ulcerative colitis [UC]. Unlike endoscopy, where standardized scoring for evaluation of disease activity is widely used, scoring for UC with IUS is currently unavailable. Therefore, we conducted a study to assess the reliability of IUS in UC among expert sonographists in order to identify robust parameters. METHODS: Thirty patients with both clinically active [25] and quiescent [five] UC were included. Six expert sonographers first agreed upon key IUS parameters and grading, including bowel wall thickness [BWT], colour Doppler signal [CDS], inflammatory fat [i-fat], loss of bowel wall stratification [BWS], loss of haustrations and presence of lymph nodes. Thirty video-recorded cases were blindly reviewed. RESULTS: Inter-observer agreement was almost perfect for BWT (intra-class correlation coefficient [ICC]: 0.96) and substantial for CDS [κ = 0.63]. Agreement was moderate for presence of lymph nodes [κ = 0.41] and fair for presence of i-fat [κ = 0.36], BWS [κ = 0.24] and loss of haustrations [κ = 0.26]. Furthermore, there was substantial agreement for presence of disease activity on IUS [κ = 0.77] and almost perfect agreement for disease severity [ICC: 0.93]. Most individual parameters showed a strong association with IUS disease activity as measured by the six readers. CONCLUSION: IUS is a reliable imaging modality to assess disease activity and severity in UC. Important individual parameters such as BWT and CDS are reliable and could be incorporated in a future UC scoring index. Standardized acquisition and assessment of UC utilizing IUS with established reliability is important to expand the use of IUS globally.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.080
metaresearch head score (Gemma)0.140
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: Observational
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.080
Threshold uncertainty score0.421

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0800.140
Meta-epidemiology (narrow)0.0010.001
Meta-epidemiology (broad)0.0010.001
Bibliometrics0.0020.001
Science and technology studies0.0010.002
Scholarly communication0.0010.002
Open science0.0010.003
Research integrity0.0010.001
Insufficient payload (model declined to judge)0.0010.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.026
GPT teacher head0.285
Teacher spread0.259 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designObservational
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations92
Published2020
Admission routes1
Has abstractyes

Explore more

Same venueJournal of Crohn s and ColitisSame topicInflammatory Bowel DiseaseFrench-language works237,207