P119 There is minimal agreement on the recognition of deep ulcers as seen on endoscopy in patients with Inflammatory Bowel Disease: a national survey
Bibliographic record
Abstract
Abstract Background Deep ulcers have been described as a marker of severe disease phenotype in patients with IBD and play a role in choice or escalation of therapy. However, there is no agreed upon characterization of deep ulcers in the literature. We therefore assessed Canadian gastroenterologists’ ability to identify the presence of deep ulcers on endoscopic images. Methods We present a post-hoc analysis of a cross-sectional questionnaire from gastroenterologists across Canada (March-October 2017). Three IBD experts independently rated 20 ileocolonoscopy images of single bowel segments. They described the images by selecting descriptors from a list developed a priori. Images described by all 3 experts as having “deep ulcers” were retained for analysis (5 images). Survey participants similarly applied descriptors from the same list to each endoscopy image. We examined the percent agreement between the gastroenterologists and the experts. The percent agreement for each question was summarized using median and IQR. Difference in median scores in physician subgroups was determined using the Mann-Whitney U test. We also assessed the inter-observer agreement on the presence of deep ulcers amongst gastroenterologists using Fleiss Kappa. Results 131 gastroenterologists participated in the study. The majority (55.7%) were between 36–50 years old. 48% were in practice for less than 10 years. 59.5% practiced in an academic setting. 9.9% of responders were pediatric gastroenterologists. The median agreement between the gastroenterologists and the experts was 30.5% (30.5–76.3), indicating an under-recognition of deep ulcers. As a group, inter-observer agreement on the presence of deep ulcers was minimal (k = 0.39, CI: 0.16–0.65). Inter-observer agreement was lower in the academic setting than the community setting (k= 0.39 vs k= 0.49 respectively) and was fairly similar in those with more experience compared to those with less experience (k= 0.40 for less than 10 years in practice, vs k= 0.36 for those with greater than 10 years in practice). Compared to the experts’ responses, there was no significant difference if the physician practiced in an academic vs community setting (49.24% correct identification vs 54.52%, p = 0.6004) or if the physician had less than 10 years’ experience vs greater than 10 years’ experience (48.26% vs 53.44%, p =0.2948). Conclusion Deep ulcers have been described as a critical indicator of disease severity in patients with IBD. However, our study shows that there is poor agreement between physicians in identifying this important feature. This indicates the need for a standardized definition of deep ulcers to prevent undertreatment of patients who require escalated therapy.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.015 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.002 | 0.003 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.004 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".