A234 INTER AND INTRA RATER RELIABILITY OF PHOTO DOCUMENTATION IN COLONOSCOPY
Bibliographic record
Abstract
Colonoscopy is one of the most powerful methods employed for early detection and prevention of colorectal cancer. Visualization of the cecum is significant as it indicates that a complete colonoscopy has been performed. Photo-documentation of the ileocecal valve has been utilized to document a complete colonoscopy. The effectiveness of this guideline as a screening tool used by endoscopists in the Saskatoon Health Region remains to be researched. To determine the inter- and intra-rater reliability amongst endoscopists in regards to colonoscopy photodocumentation. The endoscopic-photo at the moment when the endoscopist believes they have completed the colonoscopy (reached the ileocecal valve) was recorded. As well, endoscopists were instructed to get one photograph of the colon, not of the ileocecal valve but similar in appearance, for every 4 positive scopes. A total of 20 pictures were collected from 3 endoscopists and 10 pictures from 2 endoscopists. After the completion of all 80 colonoscopies, the photographs were compiled, randomized and presented to each of the 5 endoscopists. The endoscopists were required review all the 80 photographs, within a standardized time period, and grade whether the photo-documentation indicated that cecum had been reached or not. The data was recorded and a correlation coefficient was performed to analyze the intra-rater and inter-rater reliability. Overall agreement between stated image region that was taken and physician blinded review occurred in 365 of 400 reviews (91%). Among those images taken at the cecum/appendix, 297/320 reviews (93%) labeled these images as being from this region (sensitivity). Among those images not from the cecum/appendix, 68/80 reviews (85%) recognized these images as not begin from this region (specificity). As such 12/80 reviews (15%) of images taken prior to the cecum/appendix were mistakenly thought to be from the cecum/appendix (i.e. false positive). Overall positive and negative predictive values were 96% and 75% respectively. Overall agreement across multiple raters using Fleiss’s Kappa statistic = 0.69, suggesting substantial agreement beyond that occurring by chance. Of the 80 images assessed, 6 (8%) were not rated the same way when reassessed by the endoscopist who originally took the image. Three were falsely labeled as negative and three were falsely labeled as positive. Documenting colonoscopy with static images of the ileocecal valve is only partially successful as an objective measure if a total colonoscopy has been performed at the Saskatoon Health Region. Authors believe still photography of the ileocecal valve may not be convincing in all cases. A combination of photographs (involving the ileocecal valve and appendiceal orifice) or documentation of the terminal ileum could be more efficacious as an objective measure of a total colonoscopy. None
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.118 | 0.166 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.003 | 0.002 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".