The American Society of Colon and Rectal Surgeons Assessment Tool for Performance of Laparoscopic Colectomy
Bibliographic record
Abstract
BACKGROUND: The lack of consensus for performance assessment of laparoscopic colorectal resection is a major impediment to quality improvement. OBJECTIVE: The purpose of this study was to develop and assess the validity of an evaluation tool for laparoscopic colectomy that is feasible for wide implementation. DESIGN: During the pilot phase, a small group of experts modified previous assessment tools by watching videos for laparoscopic right colectomy with the following categories of experience: novice (less than 20 cases), intermediate (50-100 cases), and expert (more than 500 cases). After achieving sufficient reliability (κ > 0.8), a user-friendly tool was validated among a large group of blinded, trained experts. SETTING: The study was conducted through the American Society of Colon and Rectal Surgeons Operative Competency Evaluation Committee. PATIENTS: Raters were from the Operative Competency Evaluation Committee of the American Society of Colon and Rectal Surgeons. MAIN OUTCOME MEASURES: Assessment tool reliability and internal consistency were measured. RESULTS: From October 2014 through February 2015, 4 groups of 5 raters blinded to surgeon skill level evaluated 6 different laparoscopic right colectomy videos (novice = 2, intermediate = 2, expert = 2). The overall Cronbach α was 0.98 (>0.9 = excellent internal consistency). The intraclass correlation for the overall assessment was 0.93 (range, 0.77-0.93) and was >0.74 (excellent) for each step. The average scores (scale, 1-5) for experts were significantly better than those in the intermediate category, with a mean (SD) of 4.51 (0.56) versus 2.94 (0.56; p = 0.003). Videos in the intermediate group scored more favorably than beginner videos for each individual step and overall performance (mean (SD) = 3.00 (0.32) vs 1.78 (0.42); p = 0.006). LIMITATIONS: The study was limited by rater bias to technique and style. CONCLUSIONS: The unique and robust methodology in this trial produced an assessment tool that was feasible for raters to use when assessing videotaped laparoscopic right hemicolectomies. The potential applications for this new tool are widespread, including both training and evaluation of competence at the attending level. See Video Abstract at http://links.lww.com/DCR/A369, http://links.lww.com/DCR/A370, http://links.lww.com/DCR/A371.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.008 | 0.027 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.004 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".