Reliability of the Modified Clavien-Dindo-Sink Complication Classification System in Pediatric Orthopaedic Surgery
Bibliographic record
Abstract
BACKGROUND: are commonly used. The Clavien-Dindo-Sink complication classification system has demonstrated high interrater and intrarater reliability for hip-preservation surgery and has increasingly been used within other orthopaedic subspecialties. This classification system is based on the magnitude of treatment required and the potential for each complication to result in long-term morbidity. The purpose of the current study was to modify the Clavien-Dindo-Sink system for application to all orthopaedic procedures (including those involving the spine and the upper and lower extremity) and to determine interrater and intrarater reliability of this modified system in pediatric orthopaedic surgery cases. METHODS: The Clavien-Dindo-Sink complication classification system was modified for use with general orthopaedic procedures. Forty-five pediatric orthopaedic surgical scenarios were presented to 7 local fellowship-trained pediatric orthopaedic surgeons at 1 center to test internal reliability, and 48 scenarios were then presented to 15 pediatric orthopaedic surgeons across the United States and Canada to test external reliability. Surgeons were trained to use the system and graded the scenarios in a random order on 2 occasions. Fleiss and Cohen kappa (κ) statistics were used to determine interrater and intrarater reliabilities, respectively. RESULTS: The Fleiss κ value for interrater reliability (and standard error) was 0.76 ± 0.01 (p < 0.0001) and 0.74 ± 0.01 (p < 0.0001) for the internal and external groups, respectively. For each grade, interrater reliability was good to excellent for both groups, with an overall range of 0.53 for Grade I to 1 for Grade V. The Cohen κ value for intrarater reliability was excellent for both groups, ranging from 0.83 (95% confidence interval [CI], 0.71 to 0.95) to 0.98 (95% CI, 0.94 to 1.00) for the internal test group and from 0.83 (95% CI, 0.73 to 0.93) to 0.99 (95% CI, 0.97 to 1.00) for the external test group. CONCLUSIONS: The modified Clavien-Dindo-Sink classification system has good interrater and excellent intrarater reliability for the evaluation of complications following pediatric orthopaedic upper extremity, lower extremity, and spine surgery. Adoption of this reproducible, reliable system as a standard of reporting complications in pediatric orthopaedic surgery, and other orthopaedic subspecialties, could be a valuable tool for improving surgical practices and patient outcomes.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.019 | 0.082 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".