A Reliability Study: Strong Inter-Observer Agreement of an Expert Panel for Intestinal Ultrasound in Ulcerative Colitis
Bibliographic record
Abstract
BACKGROUND: Intestinal ultrasound [IUS] is a promising and non-invasive cross-sectional imaging modality in the diagnosis and monitoring of ulcerative colitis [UC]. Unlike endoscopy, where standardized scoring for evaluation of disease activity is widely used, scoring for UC with IUS is currently unavailable. Therefore, we conducted a study to assess the reliability of IUS in UC among expert sonographists in order to identify robust parameters. METHODS: Thirty patients with both clinically active [25] and quiescent [five] UC were included. Six expert sonographers first agreed upon key IUS parameters and grading, including bowel wall thickness [BWT], colour Doppler signal [CDS], inflammatory fat [i-fat], loss of bowel wall stratification [BWS], loss of haustrations and presence of lymph nodes. Thirty video-recorded cases were blindly reviewed. RESULTS: Inter-observer agreement was almost perfect for BWT (intra-class correlation coefficient [ICC]: 0.96) and substantial for CDS [κ = 0.63]. Agreement was moderate for presence of lymph nodes [κ = 0.41] and fair for presence of i-fat [κ = 0.36], BWS [κ = 0.24] and loss of haustrations [κ = 0.26]. Furthermore, there was substantial agreement for presence of disease activity on IUS [κ = 0.77] and almost perfect agreement for disease severity [ICC: 0.93]. Most individual parameters showed a strong association with IUS disease activity as measured by the six readers. CONCLUSION: IUS is a reliable imaging modality to assess disease activity and severity in UC. Important individual parameters such as BWT and CDS are reliable and could be incorporated in a future UC scoring index. Standardized acquisition and assessment of UC utilizing IUS with established reliability is important to expand the use of IUS globally.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".