The ACS NSQIP Risk Calculator Is a Fair Predictor of Acute Periprosthetic Joint Infection
Bibliographic record
Abstract
BACKGROUND: Periprosthetic joint infection (PJI) is a severe complication from the patient's perspective and an expensive one in a value-driven healthcare model. Risk stratification can help identify those patients who may have risk factors for complications that can be mitigated in advance of elective surgery. Although numerous surgical risk calculators have been created, their accuracy in predicting outcomes, specifically PJI, has not been tested. QUESTIONS/PURPOSES: (1) How accurate is the American College of Surgeons National Surgical Quality Improvement Program (ACS NSQIP) Surgical Site Infection Calculator in predicting 30-day postoperative infection? (2) How accurate is the calculator in predicting 90-day postoperative infection? METHODS: We isolated 1536 patients who underwent 1620 primary THAs and TKAs at our institution during 2011 to 2013. Minimum followup was 90 days. The ACS NSQIP Surgical Risk Calculator was assessed in its ability to predict acute PJI within 30 and 90 days postoperatively. Patients who underwent a repeat surgical procedure within 90 days of the index arthroplasty and in whom at least one positive intraoperative culture was obtained at time of reoperation were considered to have PJI. A total of 19 cases of PJI were identified, including 11 at 30 days and an additional eight instances by 90 days postoperatively. Patient-specific risk probabilities for PJI based on demographics and comorbidities were recorded from the ACS NSQIP Surgical Risk Calculator website. The area under the curve (AUC) for receiver operating characteristic (ROC) curves was calculated to determine the predictability of the risk probability for PJI. The AUC is an effective method for quantifying the discriminatory capacity of a diagnostic test to correctly classify patients with and without infection in which it is defined as excellent (AUC 0.9-1), good (AUC 0.8-0.89), fair (AUC 0.7-0.79), poor (AUC 0.6-0.69), or fail/no discriminatory capacity (AUC 0.5-0.59). A p value of < 0.05 was considered to be statistically significant. RESULTS: The ACS NSQIP Surgical Risk Calculator showed only fair accuracy in predicting 30-day PJI (AUC: 74.3% [confidence interval {CI}, 59.6%-89.0%]. For 90-day PJI, the risk calculator was also only fair in accuracy (AUC: 71.3% [CI, 59.9%-82.6%]). Conclusions The ACS NSQIP Surgical Risk Calculator is a fair predictor of acute PJI at the 30- and 90-day intervals after primary THA and TKA. Practitioners should exercise caution in using this tool as a predictive aid for PJI, because it demonstrates only fair value in this application. Existing predictive tools for PJI could potentially be made more robust by incorporating preoperative risk factors and including operative and early postoperative variables. LEVEL OF EVIDENCE: Level III, diagnostic study.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".