Colonoscopy quality improvement after initial training: A cross-sectional study of intensive short-term training
Bibliographic record
Abstract
Abstract Background and study aims High-quality is crucial for the effectiveness of colonoscopy and can be achieved by high-quality training and verified with assessment of key performance indicators (KPIs) for colonoscopy such as cecum intubation rate (CIR), adenoma detection rate (ADR) and adequate polyp resection. Typically, trainees achieve adequate CIR after 275 procedures, but little is known about learning curves for KPIs after initial training. Methods This cross-sectional study includes work-up colonoscopies after a positive screening test with fecal occult blood testing (FIT) or sigmoidoscopy, performed by either trainees after 300 training colonoscopies or by consultants. Outcome measures were KPIs. We assessed inter-endoscopist variation in trainees and learning curves for trainees as a group. We also compared KPIs for trainees and consultants as a group. Results Data from 6,655 colonoscopies performed by 21 trainees and 921 colonoscopies performed by 17 consultants were included. Most trainees achieved target standards for main KPIs. With time, trainees shortened cecum intubation time and withdrawal time without decreasing their ADR, reduced the proportion of painful colonoscopies, and increased the adequate polyp resection rate (all P < 0.01). Compared to consultants, trainees had higher CIR (97.7 % vs. 96.3 %, P = 0.02), ADR after positive FIT (57.6 % vs. 50.3 %, P < 0.01), and proximal ADR after sigmoidoscopy screening (41.1 % vs. 29.8 %; P < 0.01), higher adequate polyp resection rate (94.9 % vs. 93.1 %, P = 0.01) and fewer serious adverse events (0.65 % vs. 1.41 %, P = 0.02). Conclusions Trainees performed high-quality colonoscopies and achieved international target standards. Several KPIs continuously improved after initial training. Trainees outperformed consultants on several KPIs.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.010 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".