A150 MODERATE AGREEMENT IN ENDOSCOPIC DISEASE SCORING OF PEDIATRIC EOSINOPHILIC ESOPHAGITIS AMONG PEDIATRIC GASTROENTEROLOGISTS IN CANADA
Bibliographic record
Abstract
Abstract Background Endoscopy is an important tool in assessing the severity of gastrointestinal diseases including Eosinophilic Esophagitis (EoE). Agreement regarding endoscopy outcomes is important when using tools such as the Endoscopic Reference Score for EoE (EREFS). Purpose Our goal was to determine interrater and intrarater agreement of EREFS among Canadian pediatric gastroenterologists. Method Survey-based study of interrater and intrarater reliability amongst pediatric gastroenterologists with interest in pediatric EoE. Participants were sourced from the Canadian Pediatric EoE Network. Participants were asked how many years of training they’ve had with endoscopy for pediatric EoE and their comfort in disease scoring for pediatric EoE. Pediatric EoE cases were identified from the pediatric EoE registry at the Stollery Children’s Hospital with an endoscopic video associated with each case. Participants were asked to score each video using the EREFS questionnaire for the proximal, middle and distal segments of the esophagus. 15 endoscopic videos were evaluated, with 3 cases provided each week over a period of 5 weeks. Additional data included ratings of the video quality and endoscopy quality. Of 15 cases, 12 were unique cases, distributed evenly in severity between no active disease to severe disease. 3 cases were repeated to assess intrarater reliability. The maximum grade of the proximal, middle and distal segments of the esophagus for each component endoscopic finding (edema, rings, exudates, furrows, strictures) were used for reliability calculations. Fleiss Kappa was calculated for all EREFS items and for each component endoscopic finding. Cohen’s Kappa was calculated to assess intrarater reliability. Result(s) Fifteen participants were recruited for the study. The participants had a median of 12 years (IQR: 7, 19) of clinical experience in endoscopy for pediatric EoE. The majority of participants were “comfortable” (i.e., 4 on 5-point scale) with EREFS scoring for pediatric EoE. Fleiss Kappa for all EREFS items was 0.481. For each component endoscopic finding (edema, rings, exudates, furrows, strictures), Fleiss Kappa was 0.365, 0.293, 0.548, 0.263, 0.445 respectively. Cohen’s Kappa had a median of 0.620 (IQR: 0.593, 0.704). The majority of raters rated video quality and endoscopy quality as “good” (i.e., 4 on 5-point scale). Conclusion(s) There is moderate interrater reliability in EREFS scoring for pediatric EoE. Interrater reliability was between fair to moderate for each component endoscopic finding. Intrarater reliability was good. This study shows there is room for improvement in disease scoring for pediatric EoE. This could be in the form of additional training, expert-defined conventions, or centralized reading which have reduced variability in endoscopic reporting for adult GI disease in past studies and could be used in a follow-up study to attempt to improve agreement. Additionally, incorporating EREFS into routine clinical practice may increase agreement amongst endoscopists. Please acknowledge all funding agencies by checking the applicable boxes below None Disclosure of Interest None Declared
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.010 | 0.028 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.002 | 0.003 |
| Science and technology studies | 0.003 | 0.002 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".