Neural Networks Accurately Predict Precise Metrics of Hospital Resource Utilization for Total Hip Arthroplasty: A Retrospective Database Study
Bibliographic record
Abstract
Abstract Background Total hip and knee arthroplasties (THAs and TKAs) are some of the most common and successful surgeries. Predicting their duration of surgery (DOS) and length of stay (LOS) has massive implications for costs and resource management. The purpose of this study was to predict the DOS and LOS of THAs using machine learning models (MLMs) based on preoperative factors. Methods The American College of Surgeons (ACS) National Surgical and Quality Improvement (NSQIP) database was queried for elective unilateral THA procedures. Multiple MLMs were constructed to predict DOS and LOS. Models were evaluated according to mean squared error (MSE), buffer accuracy, and classification accuracy. To ensure useful predictions, the results of the models were compared to a mean regressor and previous MLM predictions for primary TKAs. Results 196,942 patients were included. The neural network had the best MSE, buffer and training accuracies for both DOS and LOS. For DOS testing, the neural network MSE was 0.916, with the 30-minute buffer and ≤120 min, >120 min accuracies being 75.4% and 88.5%. For LOS testing, the neural network MSE was 0.567, with the 1-day buffer and ≤2 days, >2 accuracies being 70.3% and 80.9%. Slightly reduced performance was found for THA compared to TKA for DOS and LOS (3 to 5%), with similar important features identified. Conclusion MLMs based on preoperative factors successfully predicted the DOS and LOS of elective unilateral THAs, with similar performance to TKA. Future work should include operational factors to apply these models to real world resource optimization.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".