Length of Stay Prediction Models for Oral Cancer Surgery: Machine Learning, Statistical and <scp>ACS‐NSQIP</scp>
Bibliographic record
Abstract
OBJECTIVE: Accurate prediction of hospital length of stay (LOS) following surgical management of oral cavity cancer (OCC) may be associated with improved patient counseling, hospital resource utilization and cost. The objective of this study was to compare the performance of statistical models, a machine learning (ML) model, and The American College of Surgeons National Surgical Quality Improvement Program's (ACS-NSQIP) calculator in predicting LOS following surgery for OCC. MATERIALS AND METHODS: A retrospective multicenter database study was performed at two major academic head and neck cancer centers. Patients with OCC who underwent major free flap reconstructive surgery between January 2008 and June 2019 surgery were selected. Data were pooled and split into training and validation datasets. Statistical and ML models were developed, and performance was evaluated by comparing predicted and actual LOS using correlation coefficient values and percent accuracy. RESULTS: Totally 837 patients were selected with mean patient age being 62.5 ± 11.7 [SD] years and 67% being male. The ML model demonstrated the best accuracy (validation correlation 0.48, 4-day accuracy 70%), compared with the statistical models: multivariate analysis (0.45, 67%) and least absolute shrinkage and selection operator (0.42, 70%). All were superior to the ACS-NSQIP calculator's performance (0.23, 59%). CONCLUSION: We developed statistical and ML models that predicted LOS following major free flap reconstructive surgery for OCC. Our models demonstrated superior predictive performance to the ACS-NSQIP calculator. The ML model identified several novel predictors of LOS. These models must be validated in other institutions before being used in clinical practice. LEVEL OF EVIDENCE: 3 Laryngoscope, 134:3664-3672, 2024.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".