A272 THE CANADIAN DIRECT OBSERVATION OF PROCEDURAL SKILLS (CANDOPS) TOOL FOR ENDOSCOPIC RETROGRADE CHOLANGIOPANCREATOGRAPHY: A MULTI-CENTRE PROSPECTIVE STUDY
Bibliographic record
Abstract
Abstract Background Previous studies have demonstrated that many graduating trainees may not have all of the skills needed to independently practice endoscopic retrograde cholangiopancreatography (ERCP) safely and effectively. As a part of competency-based learning curriculum development, it is essential to provide formative feedback to trainees and ensure they acquire the knowledge and skills for independent practice. Aims To assess the performance of advanced endoscopy trainees across Canada using the Canadian Direct Observation of Procedural Skills (CanDOPS) ERCP assessment tool. Procedural items evaluated include both technical (cannulation, sphincterotomy, stone extraction, tissue sampling, and stent placement) and non-technical (leadership, communication and teamwork, judgment and decision making) skills. Methods We conducted a prospective national multi-centre prospective study. Advanced endoscopy trainees with at least two years of gastroenterology training or five years of general surgery in North America and minimal experience performing ERCPs (less than 100 ERCP procedures) were invited to participate. The CanDOPS tool was used to measure every fifth ERCP performed by trainees over a 12-month fellowship training period. ERCPs were evaluated by experienced staff endoscopists at each study site under standard clinical protocol. Cumulative sum (CUSUM) analyses were used to generate learning curves. Results The data from five Canadian sites and 11 trainees participated in the study. A total of 261ERCP evaluations were completed. Median number of evaluations by site and trainee was 49 (IQR 31–76) and 15 (IQR 11–45). Median number of cases trainees performed prior to their ERCP training was 50 (IQR 25–400). There was a significant improvement in almost all scores over time, including selective cannulation, sphincterotomy, biliary stenting and all non-technical skills (P<0.01). CUSUM analyses using acceptable and unacceptable failure rates of 20% and 50% demonstrated trainees achieved competency for most measures in their final month of their training. Competency in tissue sampling was not achieved within a one-year training period. Conclusions This is the first ERCP performance evaluation tool that examines multiple technical and non-technical aspects of the procedure. Although trainee ERCP skills do improve during their training period, there exists a notable variability in time to competency for the different skills measured using the CanDOPS tool. Large prospective research is required to determine if competency is achieved using more stringent definitions of ERCP competency and to determine factors associated with reaching competency. Funding Agencies None
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.005 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.002 |
| Science and technology studies | 0.002 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".