Mucosal Features of Colonic Crohn’s Disease Determined by Confocal Laser Endomicroscopy (CLE): An Inter-observer Agreement Study
Bibliographic record
Abstract
Introduction: Crohn’s disease causes colonic mucosal changes that may alter permeability and crypt characteristics that may be assessed by in vivo microscopic imaging of the colonic mucosa using CLE. This study aimed to develop a quantitative image analysis technique to characterize these changes accurately and assess the reliability of the technique. Methods: In a pilot study of patients with active or inactive Crohn’s colitis or controls with normal mucosa, CLE images (EC3870K; Pentax, Tokyo, Japan), were obtained from inflamed and non-inflamed mucosa after IV fluorescein 10% (5 mL) in 30 patients. CLE images were analyzed (ImageJ, NIH, Bethesda, U.S.) to outline colonic crypts, determine greyscale density of crypts and intercryptal areas and to determine crypt number, size and shape. Using standardized criteria, 30 randomly selected images were evaluated independently by 2 observers (O1 & O2); Cohen’s kappa coefficient (linear weighting) was calculated to assess inter-observer agreement for the number of complete crypts per image, mean crypt diameter and grey scale ratio for crypt to intercryptal tissue (fluorescein density). Results: For the 30 images, the mean numbers of complete crypts were 5.33 [95% confidence interval (CI): 4.72 to 5.94] (O1) and 5.10 [4.46 to 5.74] (O2) - kappa 0.844 [0.735 to 0.953], the mean crypt diameters were 104.0 [97.6 to 110.4] (O1) and 110.6 [102.2 to 119.1] (O2) - kappa 0.694 [0.470 to 0.828] and the mean grey scale ratios were 0.67 [0.58 to 0.75] (O1) and 0.69 [0.60 to 0.78] (O2) - kappa 0.810 [0.672 to 0.948]. Bland-Altman plots showed no systematic inter-observer differences for crypt numbers or greyscale ratios (Figure 1).Figure 1Conclusion: Good inter-observer agreement was obtained for quantitative analysis, using standardized criteria, for objective characterization of CLE images obtained from inflamed and non-inflamed colonic mucosa. This technique allows for standardized evaluation of colonic CLE images to evaluate extravasation of fluorescein associated with acute inflammation and changes in crypt size and density associated with more chronic inflammation; this will allow objective quantification of inflammatory changes and the effects of treatment on disease activity. Disclosure - Dr. David Armstrong - Consultant: Abbvie, Forest, Takeda; Research: Abbvie; Speakers Bureau: Aptalis, Janssen, Shire, Takeda; Educational Events: Abbvie, Aptalis, Ferring, Janssen, Shire, Takeda. No other authors have any relevant financial relationships. This research was supported by an industry grant from AbbVie.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.035 | 0.041 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".