S550 Development of the Toronto Upper Gastrointestinal Cleaning Score: A Delphi Study
Bibliographic record
Abstract
Introduction: Esophagogastroduodenoscopy (EGD) is essential for the evaluation of the foregut and is dependent on visualization of the upper gastrointestinal (GI) mucosa pathology. Foregut evaluation can be limited by the presence of mucus, foam, bubbles, and solid materials. Inadequate visualization may necessitate repeat endoscopy, exposing the patient to additional procedural risk. Currently, there is no standardized method to assess mucosal visualization for use in clinical or research settings. By using Delphi methodology, we aimed to develop and establish the content validity of the Toronto Upper Gastrointestinal Cleaning Score (TUGCS). Methods: We invited an international panel of endoscopy experts and educators to rate potential anatomical items and their associated anchors for importance as indicators of adequacy of mucosal visualization during EGD. The survey included questions regarding the TUGCS with Likert scales, wherein each participant reported their agreement with given statements on a scale of 1 (I strongly disagree with this statement) to 5 (I strongly agree with this statement). After each round in the Delphi process, we evaluated agreement for each survey item and sent back a revised version of the TUGCS to the expert panel for further ratings until we reached a consensus. We defined consensus a priori as ≥80% of experts in a given round, scoring ≥4 on all survey items. Results: Fourteen experts agreed to be part of the Delphi panel. We initially generated an anatomical framework to represent the upper GI mucosa and anchors for each mucosal portion to represent various levels of visibility through a systematic review. After three rounds of surveys, with response rates of 100%, 100%, and 71% respectively, consensus was achieved. The final TUGCS includes four anatomical areas (fundus, body, antrum, duodenum) and mucosal visualization anchors ranging from 0 (any solid food, blood or blood clots, or other content that could not be suctioned or washed, or an obstruction that prevented adequate visualization of the majority of an anatomical area) to 3 (entire mucosa well seen without the need for suctioning or washing) (Figure 1). Conclusion: We developed and generated content validity evidence for the TUGCS through rigorous Delphi methodology, reflective of practice across different centers. We are currently gathering validity evidence for the TUGCS in the effort of creating a tool that may be used to judge mucosal visualization for EGD in research and clinical settings.Figure 1.: Toronto Upper Gastrointestinal Cleaning Score
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.045 | 0.053 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.003 | 0.002 |
| Science and technology studies | 0.002 | 0.002 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.001 | 0.004 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.005 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".