A32 DEVELOPMENT AND VALIDATION OF THE TORONTO UPPER GASTROINTESTINAL CLEANING SCORE
Bibliographic record
Abstract
Abstract Background High quality esophagogastroduodenoscopy (EGD) depends on the ability to appropriately visualize upper gastrointestinal (GI) mucosa pathology. Evaluation can be limited by the presence of mucus, foam, bubbles and solid materials. Currently, there is no standardized method to assess mucosal visualization for use in clinical or research settings. Aims To develop and establish the content validity of the Toronto Upper Gastrointestinal Cleaning Score (TUGCS) and evaluate its interrater reliability. Methods An international panel of endoscopy experts rated potential items and their associated anchors for importance as indicators of adequacy of mucosal visualization during EGD. The survey utilized a Likert scale (1 (strongly disagree) to 5 (strongly agree)). The Delphi process was repeated until consensus was reached. Consensus was defined priori as ≥80% of experts in a given round scoring ≥4 on all survey items. To assess content validity, 48 EGD procedures were evaluated in real-time by two endoscopist reviewers using the TUGCS at a single institution. The interrater agreement between assessments was calculated for TUGCS total scores using intraclass correlation coefficient, one-way random effects model (ICC 1,1). Results Fourteen experts agreed to be part of the Delphi panel. An anatomical framework representing the upper GI mucosa and anchors for each mucosal portion representing various levels of visibility was generated through systematic review. Three survey rounds, with response rates of 100%, 100% and 71% respectively, achieved consensus. The final TUGCS includes four anatomical areas (fundus, body, antrum, duodenum) and mucosal visualization anchors ranging from 0 to 3 (Figure 1). TUGCS was used to assess foregut cleaning in 48 procedures (Table 1). The mean TUGCS for staff and trainee were 8.1 (±2.4) and 8.1 (±2.6), respectively. The ICC was 0.78 (95% confidence interval 0.62–0.88) indicating good reliability. Conclusions We developed and generated content validity evidence for the TUGCS through rigorous Delphi methodology, reflective of practice across different centres. Planned as future research is a video survey distributed to endoscopists internationally to further validate the TUGCS to create a tool that may be used to judge mucosal visualization for EGD in research and clinical settings. Funding Agencies None
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.023 | 0.052 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.003 |
| Bibliometrics | 0.006 | 0.003 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.001 | 0.003 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.004 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".