Development of a Clinical Prediction Model for Anastomotic Leakage in Colorectal Surgery
Bibliographic record
Abstract
Importance: Anastomotic leakage is a severe complication after colorectal surgery that can significantly affect clinical outcomes and patient prognosis. Although multiple risk factors for anastomotic leakage have been identified, predicting the incidence of anastomotic leakage remains difficult. Objective: To develop a machine learning predictive model that combines logits from different base models to predict the incidence of anastomotic leakage more accurately. Design, Setting, and Participants: This retrospective prognostic study was conducted across 13 centers in Europe and North America between January 1, 2012, and December 31, 2020, and with data from the European Society of Coloproctology (ESCP). The study population included patients 18 years or older who underwent colon resection with anastomosis and had a follow-up of at least 6 months. Patients with metastatic disease and insufficient follow-up were excluded. A total of 6079 patients from 13 centers across different countries were included in the analysis with an additional 3041 patients from the ESCP. Data analysis was conducted from October 2024 to January 2025. Exposure: Colon anastomosis for various reasons, including neoplasia, diverticulitis, ischemia, iatrogenic or traumatic perforation, or inflammatory bowel disease. Main Outcomes and Measures: The primary outcome of this study was to predict anastomotic leakage using a cross-attention-based meta-model, which integrated predictions from base models and patient-level clinical features. The F1 score was used to assess the prediction performance of the models. Results: A total of 9120 patients (mean [SD] age, 61.26 [15.71] years; 4636 [50.8%] male) from the 513 centers and the ESCP were included in the analysis. There was no difference in distribution between the 13 centers and the ESCP dataset. The model's predictive performance achieved overall F1 scores of 87% (95% CI, 78%-95%) and 70% across tests using a cross-validation process and external validation test set, respectively. The meta-model's performance was better than that of its component parts (eg, CatBoost: cross-validation F1 score, 87% [95% CI, 78%-95%]; external validation F1 score, 70%). Conclusions and Relevance: In this prognostic study of patients who underwent colon resection with anastomosis, a meta-model was found to improve predictive accuracy. Additional prospective studies are needed to assess the clinical utility of the model in real-time settings and its integration into surgical workflows.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".