An International Expert Delphi Consensus on Defining Textbook Outcome in Liver Surgery (TOLS)
Bibliographic record
Abstract
OBJECTIVE: To reach global expert consensus on the definition of TOLS in minimally invasive and open liver resection among renowned international expert liver surgeons using a modified Delphi method. BACKGROUND: Textbook outcome is a novel composite measure combining the most desirable postoperative outcomes into one single measure and representing the ideal postoperative course. Despite a recently developed international definition of Textbook Outcome in Liver Surgery (TOLS), a standardized and expert consensus-based definition is lacking. METHODS: This international, consensus-based, qualitative study used a Delphi process to achieve consensus on the definition of TOLS. The survey comprised 6 surgical domains with a total of 26 questions on individual surgical outcome variables. The process included 4 rounds of online questionnaires. Consensus was achieved when a threshold of at least 80% agreement was reached. The results from the Delphi rounds were used to establish an international definition of TOLS. RESULTS: In total, 44 expert liver surgeons from 22 countries and all 3 major international hepato-pancreato-biliary associations completed round 1. Forty-two (96%), 41 (98%), and 41 (98%) of the experts participated in round 2, 3, and 4, respectively. The TOLS definition derived from the consensus process included the absence of intraoperative grade ≥2 incidents, postoperative bile leakage grade B/C, postoperative liver failure grade B/C, 90-day major postoperative complications, 90-day readmission due to surgery-related major complications, 90-day/in-hospital mortality, and the presence of R0 resection margin. CONCLUSIONS: This is the first study providing an international expert consensus-based definition of TOLS for minimally invasive and open liver resections by the use of a formal Delphi consensus approach. TOLS may be useful in assessing patient-level hospital performance and carrying out international comparisons between centers with different clinical practices to further improve patient outcomes.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".