MétaCan
Menu
Back to cohort
Record W4294990764 · doi:10.1093/bjs/znac300

Benchmarks in colorectal surgery: multinational study to define quality thresholds in high and low anterior resection

2022· article· en· W4294990764 on OpenAlexaff
Roxane D. Staiger, Fabian Rössler, Min Jung Kim, Carl J. Brown, Loris Trenti, Takeshi Sasaki, Deniz Uluk, Juan Pablo Campana, Massimo Giacca, Boris Schiltz, Renu R. Bahadoer, Kai-Yin Lee, Bruna Elisa Catin Kupper, Katherine Y. Hu, Francesco Corcione, Steven Ronald Paredes, Sebastiano Spampati, Kristjan Ukegjini, Bartlomiej Jedrzejczak, Daniël Langer, Áine Stakelum, Ji Won Park, P. Terry Phang, Sebastiano Biondo, Masaaki Ito, Felix Aigner, Carlos Vaccaro, Yves Panís, A. Kartheuser, Koen Peeters, Ker‐Kan Tan, Samuel Aguiar, Kirk Ludwig, Christopher J. Young, Adam Dziki, Miroslav Ryska, D. C. Winter, John T. Jenkins, Robin H. Kennedy, Pierre‐Alain Clavien, Milo A. Puhan, Matthias Turina

Bibliographic record

VenueBritish journal of surgery · 2022
Typearticle
Languageen
FieldMedicine
TopicColorectal Cancer Surgical Treatments
Canadian institutionsSt. Paul's HospitalUniversity of British Columbia
Fundersnot available
KeywordsMedicinePercentileRetrospective cohort studySurgeryColorectal cancerBenchmarkingQuality managementBenchmark (surveying)CancerInternal medicineStatisticsOperations management

Abstract

fetched live from OpenAlex

BACKGROUND: Benchmark comparisons in surgery allow identification of gaps in the quality of care provided. The aim of this study was to determine quality thresholds for high (HAR) and low (LAR) anterior resections in colorectal cancer surgery by applying the concept of benchmarking. METHODS: This 5-year multinational retrospective study included patients who underwent anterior resection for cancer in 19 high-volume centres on five continents. Benchmarks were defined for 11 relevant postoperative variables at discharge, 3 months, and 6 months (for LAR). Benchmarks were calculated for two separate cohorts: patients without (ideal) and those with (non-ideal) outcome-relevant co-morbidities. Benchmark cut-offs were defined as the 75th percentile of each centre's median value. RESULTS: A total of 3903 patients who underwent HAR and 3726 who had LAR for cancer were analysed. After 3 months' follow-up, the mortality benchmark in HAR for ideal and non-ideal patients was 0.0 versus 3.0 per cent, and in LAR it was 0.0 versus 2.2 per cent. Benchmark results for anastomotic leakage were 5.0 versus 6.9 per cent for HAR, and 13.6 versus 11.8 per cent for LAR. The overall morbidity benchmark in HAR was a Comprehensive Complication Index (CCI®) score of 8.6 versus 14.7, and that for LAR was CCI® score 11.9 versus 18.3. CONCLUSION: Regular comparison of individual-surgeon or -unit outcome data against benchmark thresholds may identify gaps in care quality that can improve patient outcome.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.004
metaresearch head score (Gemma)0.001
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: Observational
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.028
Threshold uncertainty score0.611

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0040.001
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0010.000
Bibliometrics0.0010.001
Science and technology studies0.0000.000
Scholarly communication0.0000.000
Open science0.0000.000
Research integrity0.0000.000
Insufficient payload (model declined to judge)0.0000.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.034
GPT teacher head0.312
Teacher spread0.278 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designObservational
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations24
Published2022
Admission routes1
Has abstractyes

Explore more

Same venueBritish journal of surgerySame topicColorectal Cancer Surgical TreatmentsFrench-language works237,207