A scoping review of complication prediction models in spinal surgery: An analysis of model development, validation and impact
Bibliographic record
Abstract
Background: Predictive analytics are being used increasingly in the field of spinal surgery with the development of models to predict post-surgical complications. Predictive models should be valid, generalizable, and clinically useful. The purpose of this review was to identify existing post-surgical complication prediction models for spinal surgery and to determine if these models are being adequately investigated with internal/external validation, model updating and model impact studies. Methods: This was a scoping review of studies pertaining to models for the prediction of post-surgical complication after spinal surgery published over 10 years (2010-2020). Qualitative data was extracted from the studies to include study classification, adherence to Transparent Reporting of a multivariable prediction model for Individual Prognosis Or Diagnosis (TRIPOD) guidelines and risk of bias (ROB) assessment using the Prediction model study Risk Of Bias Assessment Tool (PROBAST). Model evaluation was determined using area under the curve (AUC) when available. The Preferred Reporting Items for Systematic Reviews and Meta-Analysis (PRISMA) statement was used as a basis for the search methodology in four different databases. Results: Thirty studies were included in the scoping review and 80% (24/30) included model development with or without internal validation. Twenty percent (6/30) were exclusively external validation studies and only one study included an impact analysis in addition to model development and internal validation. Two studies referenced the TRIPOD guidelines and there was a high ROB in 100% of the studies using the PROBAST tool. Conclusions: The majority of post-surgical complication prediction models in spinal surgery have not undergone standardized model development and internal validation or adequate external validation and impact evaluation. As such there is uncertainty as to their validity, generalizability, and clinical utility. Future efforts should be made to use existing tools to ensure standardization in development and rigorous evaluation of prediction models in spinal surgery.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.170 | 0.501 |
| Meta-epidemiology (narrow) | 0.003 | 0.003 |
| Meta-epidemiology (broad) | 0.011 | 0.023 |
| Bibliometrics | 0.039 | 0.034 |
| Science and technology studies | 0.002 | 0.003 |
| Scholarly communication | 0.009 | 0.008 |
| Open science | 0.005 | 0.006 |
| Research integrity | 0.004 | 0.004 |
| Insufficient payload (model declined to judge) | 0.004 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".